Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Epitope mapping of a monoclonal antibody against human thrombin by H/D-exchange mass spectrometry reveals selection of a diverse sequence in a highly conserved protein.

The epitope of a monoclonal antibody raised against human thrombin has been determined by hydrogen/deuterium exchange coupled to MALDI mass spectrometry. The antibody epitope was identified as the surface of thrombin that retained deuterium in the presence of the monoclonal antibody compared to control experiments in its absence. Covalent attachment of the antibody to protein G beads and efficient elution of the antigen after deuterium exchange afforded the analysis of all possible epitopes in a single MALDI mass spectrum. The epitope, which was discontinuous, consisting of two peptides close to anion-binding exosite I, was readily identified. The epitope overlapped with, but was not identical to, the thrombomodulin binding site, consistent with inhibition studies. The antibody bound specifically to human thrombin and not to murine or bovine thrombin, although these proteins share 86% identity with the human protein. Interestingly, the epitope turned out to be the more structured of two surface regions in which higher sequence variation between the three species is seen.

Amino Acid Sequence↗

Diversity of nucleotide sequences in hypervariable region 1 of hepatitis C virus in Japanese patients with chronic hepatitis C of unknown mode of transmission.

We evaluated the sequence diversity of the hypervariable region 1 (HVR1) of hepatitis C virus (HCV) in HCV-infected patients in whom the mode of transmission is unknown. The sequence diversity of HVR1 in 26 Japanese patients with chronic HCV infection of unknown mode of transmission (UT) was compared with 17 patients with chronic posttransfusion hepatitis C in whom only a single HCV infection had occurred (PH), and with 18 patients with hemophilia with chronic HCV infection who might have been multiply infected with HCV (HE). The diversity of HVR1 was evaluated by direct sequencing after PCR amplification of HVR1. The sequence diversity of HVR1 was 10.1 +/- 7.7% in UT, 2.7 +/- 2.8% in PH, and 14.6 +/- 6.9% in HE. The diversity of the patients with unknown transmission was greater than that of the posttransfusion hepatitis patients, and in some patients it was similar to that of multitransfused hemophiliac patients (UT vs PH, P = 0.0004; UT vs HE, P = 0.04; and HE vs PH, P < 0.0001). Multiple infections with HCV could have occurred frequently in patients with chronic HCV infection in whom the mode of transmission was unknown, which increased the sequence diversity of HVR1 of these patients.

Adult↗

Identification and properties of type I-signal peptidases of Bacillus amyloliquefaciens.

The use of Bacillus amyloliquefaciens for enzyme production and its exceptional high protein export capacity initiated this study where the presence and function of multiple type I signal peptidase isoforms was investigated. In addition to type I signal peptidases SipS(ba) [Meijer, W.J.J., de Jong, A., Bea, G., Wisman, A., Tjalsma, H., Venema, G., Bron, S. & van Dijl, J.M. (1995) Mol. Microbiol. 17, 621-631] and SipT(ba) [Hoang, V. & Hofemeister, J. (1995) Biochim. Biophys. Acta 1269, 64-68] which were previously identified, here we present evidence for two other Sip-like genes in B. amyloliquefaciens. Same map positions as well as sequence motifs verified that these genes encode homologues of Bacillus subtilis SipV and SipW. SipU-encoding DNA was not found in B. amyloliquefaciens. SipW-encoding DNA was also found for other Bacillus strains representing different phylogenetic groups, but not for Bacillus stearothermophilus and Thermoactinomyces vulgaris. The absence of these genes, however, could have been overlooked due to sequence diversity. Sequence alignments of 23 known Sip-like proteins from Bacillus origin indicated further branching of the P-group signal peptidases into clusters represented by B. subtilis SipV, SipS-SipT-SipU and B. anthracis Sip3-Sip5 proteins, respectively. Each B. amyloliquefaciens sip(ba) gene was expressed in an Escherichia coli LepBts mutant and tested for genetic complementation of the temperature sensitive (TS) phenotype as well as pre-OmpA processing. Although SipS(ba) as well as SipT(ba) efficiently restored processing of pre-OmpA in E. coli, only SipS(ba) supported growth at TS conditions, indicating functional diversity. Changed properties of the sip(ba) gene disruption mutants, including cell autolysis, motility, sporulation, and nuclease activities, seemed to correlate with specificities and/or localization of B. amyloliquefaciens SipS, SipT and SipV isoforms.

Bacillus↗

Worldwide haplotype diversity and coding sequence variation at human bitter taste receptor loci.

Bitter taste perception in humans is mediated by receptors encoded by 25 genes that together comprise the TAS2R (or T2R) gene family. The ability to identify the ligand(s) for each of these receptors is dependent on understanding allelic variation in TAS2R genes, which may have a significant effect on ligand recognition. To investigate the extent of coding variation among TAS2R alleles, we performed a comprehensive evaluation of sequence and haplotype variation in the human bitter taste receptor gene repertoire. We found that these genes exhibit substantial coding sequence diversity. In a worldwide population sample of 55 individuals, we found an average of 4.2 variant amino acid positions per gene. In aggregate, the 24 genes analyzed here, along with the phenylthiocarbamide (PTC) receptor gene analyzed previously, specify 151 different protein coding haplotypes. Analyses of the ratio of synonymous and nonsynonymous nucleotide substitutions using the Ka/Ks ratio revealed an excess of amino acid substitutions relative to most other genes examined to date (Ka/Ks = 0.94). In addition, comparisons with more than 1,500 other genes revealed that levels of diversity in the TAS2R genes were significantly greater than expected (pi = 0.11%; p < 0.01), as were levels of differentiation among continental populations (FST = 0.22; p < 0.05). These diversity patterns indicate that unusually high levels of allelic variation are found within TAS2R loci and that human populations differ appreciably with respect to TAS2R allele frequencies. Diversity in the TAS2R genes may be accounted for by natural selection, which may have favored alleles responsive to toxic, bitter compounds found in plants. These findings are consistent with the view that different alleles of the TAS2R genes encode receptors that recognize different ligands, and suggest that the haplotypes we have identified will be important in studies of receptor-ligand recognition.

Alleles↗

A conserved sequence motif within the exceptionally diverse telomeric sequences of budding yeasts.

Telomeric DNA sequences have generally been found to be remarkably conserved in evolution, typically consisting of repeated, very short sequence units containing clusters of G residues. Recently however the telomeric DNA of the asexual yeast Candida albicans was shown to consist of much longer repeat units. Here we report the identification of seven additional telomeric sequences from sexual and asexual budding yeast species. The telomeric repeat units from this group of relatively closely related species show more phylogenetic diversity in length (8-25 bp), sequence, and composition than has been seen previously throughout a wide phylogenetic range of other eukaryotes. We also show that certain strains of the asexual diploid species Candida tropicalis have two forms of telomeric repeats, which appear to differ by a single base pair. Despite their great diversity, the telomeric repeat units of C. albicans, Saccharomyces cerevisiae, and all of the species we have examined in this report share a conserved approximately 6-bp motif of T and G residues resembling more typical telomeric sequences.

Base Sequence↗

A generic sequencing based typing approach for the identification of HLA-A diversity.

Sequencing Based Typing (SBT) is a generic approach for the identification of HLA-A polymorphism. This approach includes the high resolution typing of the HLA-A broad reacting groups, HLA-A subtypes and will identify new alleles directly. The SBT approach described here uses a locus specific amplification of DNA from exon 1 to exon 5. The resulting 2,022 bp PCR product serves as a template for the subsequent sequencing reactions. Amplification is followed by direct sequencing of exons 2, 3 and 4 in both orientations with fluorescently labeled primers to define all polymorphic positions leading to a high resolution typing result. In this study the sequence of exons 2 and 3 of a panel of 49 cell lines was determined. In addition, the exon 4 region of 35 cell lines was also sequenced to evaluate the exon 4 polymorphism. The HLA-A type of most of the cells could be identified by sequencing only exons 2 and 3. However, the sequence of exon 4 was required to discriminate A*0201 from A*0209 and A*0207 from A*0215N. In this panel, an identical new "HLA-A*0103" was identified in two Caucasian samples.

Alleles↗

Specific ribosomal DNA sequences from diverse environmental settings correlate with experimental contaminants.

Phylogenetic analysis of 16S ribosomal DNA (rDNA) clones obtained by PCR from uncultured bacteria inhabiting a wide range of environments has increased our knowledge of bacterial diversity. One possible problem in the assessment of bacterial diversity based on sequence information is that PCR is exquisitely sensitive to contaminating 16S rDNA. This raises the possibility that some putative environmental rRNA sequences in fact correspond to contaminant sequences. To document potential contaminants, we cloned and sequenced PCR-amplified 16S rDNA fragments obtained at low levels in the absence of added template DNA. 16S rDNA sequences closely related to the genera Duganella (formerly Zoogloea), Acinetobacter, Stenotrophomonas, Escherichia, Leptothrix, and Herbaspirillum were identified in contaminant libraries and in clone libraries from diverse, generally low-biomass habitats. The rRNA sequences detected possibly are common contaminants in reagents used to prepare genomic DNA. Consequently, their detection in processed environmental samples may not reflect environmentally relevant organisms.

Animals↗

Direct measurement of T-cell receptor repertoire diversity with AmpliCot.

Many studies require the measurement of nucleic acid sequence diversity. Here we describe a method, called AmpliCot, that measures the sequence diversity of PCR products on the basis of DNA hybridization kinetics, thereby avoiding the time, expense and biases associated with cloning and sequencing. SYBR Green dye is used to measure DNA hybridization kinetics in a homogeneous, automated fashion. PCR products are prepared in wholly double-stranded homoduplex form for a baseline measurement of DNA concentration. The DNA is melted and then reannealed under stringent conditions that allow only homoduplexes to form. The sequence diversity of a sample is proportional to the product of its concentration and the time required for it to anneal. After validating AmpliCot with a library of diverse sequences, we use it to measure the diversity of expressed rearrangements of the gene encoding the T-cell antigen receptor (TCR) beta chain. AmpliCot measurements are in good agreement with previous estimates of murine TCR repertoire diversity that required extensive cloning and sequencing.

Animals↗

Diversity of sequences of polyadenylated cytoplasmic RNA from rainbow trout (Salmo gairdnerii) testis and liver.

We have compared the sequence complexity and diversity of polyadenylated cytoplasmic RNA derived from two differentiated trout tissues: liver and testis. The kinetics of hybridization of polyadenylated RNA from each of these tissues with complementary DNA synthesized by reverse transcriptase revealed three abundance classes for liver RNA, the first comprising 4 sequences, the second 120, and the third, 20 000; in contrast, testis RNA showed only two abundance classes containing 6 and 6100 different RNA sequences, respectively and of average length 6 x 10(5) daltons. The extent of overlapping among those two RNA populations was further studied by performing heterologous annealing reactions between cDNA and a vast excess of mRNA. Liver mRNA was complementary to 80% of the testis cDNA. Conversely, testis mRNA reacted with only 25% of the liver cDNA. Experiments with fractionated cDNA probes indicated that the unshared sequences belonged mainly to the less frequent, most complex, class of mRNAs.

Animals↗

Optimal cDNA microarray design using expressed sequence tags for organisms with limited genomic information.

BACKGROUND: Expression microarrays are increasingly used to characterize environmental responses and host-parasite interactions for many different organisms. Probe selection for cDNA microarrays using expressed sequence tags (ESTs) is challenging due to high sequence redundancy and potential cross-hybridization between paralogous genes. In organisms with limited genomic information, like marine organisms, this challenge is even greater due to annotation uncertainty. No general tool is available for cDNA microarray probe selection for these organisms. Therefore, the goal of the design procedure described here is to select a subset of ESTs that will minimize sequence redundancy and characterize potential cross-hybridization while providing functionally representative probes. RESULTS: Sequence similarity between ESTs, quantified by the E-value of pair-wise alignment, was used as a surrogate for expected hybridization between corresponding sequences. Using this value as a measure of dissimilarity, sequence redundancy reduction was performed by hierarchical cluster analyses. The choice of how many microarray probes to retain was made based on an index developed for this research: a sequence diversity index (SDI) within a sequence diversity plot (SDP). This index tracked the decreasing within-cluster sequence diversity as the number of clusters increased. For a given stage in the agglomeration procedure, the EST having the highest similarity to all the other sequences within each cluster, the centroid EST, was selected as a microarray probe. A small dataset of ESTs from Atlantic white shrimp (Litopenaeus setiferus) was used to test this algorithm so that the detailed results could be examined. The functional representative level of the selected probes was quantified using Gene Ontology (GO) annotations. CONCLUSIONS: For organisms with limited genomic information, combining hierarchical clustering methods to analyze ESTs can yield an optimal cDNA microarray design. If biomarker discovery is the goal of the microarray experiments, the average linkage method is more effective, while single linkage is more suitable if identification of physiological mechanisms is more of interest. This general design procedure is not limited to designing single-species cDNA microarrays for marine organisms, and it can equally be applied to multiple-species microarrays of any organisms with limited genomic information.

Animals↗

Individual sequences in large sets of gene sequences may be distinguished efficiently by combinations of shared sub-sequences.

BACKGROUND: Most current DNA diagnostic tests for identifying organisms use specific oligonucleotide probes that are complementary in sequence to, and hence only hybridize with the DNA of one target species. By contrast, in traditional taxonomy, specimens are usually identified by 'dichotomous keys' that use combinations of characters shared by different members of the target set. Using one specific character for each target is the least efficient strategy for identification. Using combinations of shared bisectionally-distributed characters is much more efficient, and this strategy is most efficient when they separate the targets in a progressively binary way. RESULTS: We have developed a practical method for finding minimal sets of sub-sequences that identify individual sequences, and could be targeted by combinations of probes, so that the efficient strategy of traditional taxonomic identification could be used in DNA diagnosis. The sizes of minimal sub-sequence sets depended mostly on sequence diversity and sub-sequence length and interactions between these parameters. We found that 201 distinct cytochrome oxidase subunit-1 (CO1) genes from moths (Lepidoptera) were distinguished using only 15 sub-sequences 20 nucleotides long, whereas only 8-10 sub-sequences 6-10 nucleotides long were required to distinguish the CO1 genes of 92 species from the 9 largest orders of insects. CONCLUSION: The presence/absence of sub-sequences in a set of gene sequences can be used like the questions in a traditional dichotomous taxonomic key; hybridisation probes complementary to such sub-sequences should provide a very efficient means for identifying individual species, subtypes or genotypes. Sequence diversity and sub-sequence length are the major factors that determine the numbers of distinguishing sub-sequences in any set of sequences.

Algorithms↗

Phylogenetic analysis of global hepatitis E virus sequences: genetic diversity, subtypes and zoonosis.

Nucleotide sequences from a total of 421 HEV isolates were retrieved from Genbank and analysed. Phylogenetically, HEV was classified into four major genotypes. Genotype 1 was more conserved and classified into five subtypes. The number of genotype 2 sequences was limited but can be classified into two subtypes. Genotypes 3 and 4 were extremely diverse and can be subdivided into ten and seven subtypes. Geographically, genotype 1 was isolated from tropical and several subtropical countries in Asia and Africa, and genotype 2 was from Mexico, Nigeria, and Chad; whereas genotype 3 was identified almost worldwide including Asia, Europe, Oceania, North and South America. In contrast, genotype 4 was found exclusively in Asia. It is speculated that genotype 3 originated in the western hemisphere and was imported to several Asian countries such as Japan, Korea and Taiwan, while genotype 4 has been indigenous and likely restricted to Asia. Genotypes 3 and 4 were not only identified in swine but also in wild animals such as boar and a deer. Furthermore, in most areas where genotypes 3 and 4 were characterised, sequences from both humans and animals were highly conserved, indicating they originated from the same infectious sources. Based upon nucleotide differences from five phylogenies, it is proposed that five, two, ten and seven subtypes for HEV genotypes 1, 2, 3 and 4 be designated alphabetised subtypes. Accordingly, a total of 24 subtypes (1a, 1b, 1c, 1d, 1e, 2a, 2b, 3a, 3b, 3c, 3d, 3e, 3f, 3g, 3h, 3i, 3j, 4a, 4b, 4c, 4d, 4e, 4f and 4g) were given.

Animals↗

DIVAA: analysis of amino acid diversity in multiple aligned protein sequences.

MOTIVATION: Multiple alignments of proteins are an effective way of identifying conserved amino acids that provide clues to functional relationships among proteins. Quantitation of the abundances of amino acids found at each position in a sequence motif can provide a basis for understanding the structural and functional constraints at each point. Distribution of information across a motif has been used previously, but the non-intuitive nature of the analysis has limited its impact. RESULTS: Here, we introduce a quantitative measure of amino acid sequence diversity (DIVAA) that has a simple, intuitive meaning. Diversity, as a measure of sequence conservation or variation, is inextricably linked to the probability of selecting identical pairs from a distribution. We demonstrate its utility through the analysis of four populations: ATP-binding P-loops, hypervariable domains of kappa light chains, signal sequences, and the N- and C- termini of proteins. DIVAA provides a simple means to generate hypotheses concerning the contribution of individual residues to the functional and evolutionary relationships among proteins. AVAILABILITY: Access to DIVAA software is available at RELIC (http://relic.bio.anl.gov).

Algorithms↗

The human (PEDB) and mouse (mPEDB) Prostate Expression Databases.

The Prostate Expression Databases (PEDB and mPEDB) are online resources designed to allow researchers to access and analyze gene expression information derived from the human and murine prostate, respectively. Human PEDB archives more than 84 000 Expressed Sequence Tags (ESTs) from 38 prostate cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library sequence comparisons. Recent enhancements to PEDB include (i) the development of a murine prostate expression database, mPEDB, that complements the human gene expression information in PEDB, (ii) the assembly of a non-redundant sequence set or 'prostate unigene' that represents the diversity of gene expression in the prostate, and (iii) an expanded search tool that supports both text-based and BLAST queries. PEDB and mPEDB are accessible via the World Wide Web at http://www.pedb.org and http://www.mpedb.org.

Animals↗

Selection of antibody epitopes in an immunopathogenic neural autoantigen.

Results with two well-characterized self-antigens, cytochrome c and myelin basic protein, have led to differing opinions regarding the predominant specificities of autoantibodies, whether regions of sequence diversity or 'structurally inherent features' of a protein determine favored antigenic sites. To further examine this question, 16 antibody epitopes have been mapped on a highly immunopathogenic autoantigen, retinal S-antigen (S-Ag). The epitopes were characterized for: (1) sequence diversity and cross-reactivity on S-antigens from several species; (2) conformational dependency; and (3) probability of their occurrence on the surface of S-antigen. A single C-terminal region containing sequence diversity was most frequently recognized, but no evidence for recognition of any other regions of sequence diversity was found. Thirteen of 16 monoclonal antibodies raised to native S-Ag bound epitopes strongly predicted to be on the surface of S-antigen. Conversely, only one of six antibody preparations raised to peptides or affinity-purified on peptides was found to recognize an epitope predicted to be on the surface, suggesting a good correlation between specificity for conformation-dependent sites and surface probability based on the surface prediction algorithm. Three of these six antibodies which preferred denatured epitopes bound sites which overlapped or coincided with T cell sites; two of these T cell sites are immunopathogenic. The epitopes recognized on denatured antigen and peptides were similar whether the antibodies were elicited with intact human or bovine S-antigen or with cyanogen bromide-cleaved peptides. Our data suggests that in the case of S-antigen, structural features are more significant factors in epitope selection than sequence diversity.

Amino Acid Sequence↗

A substrate-specific inhibitor of protein translocation into the endoplasmic reticulum.

The segregation of secretory and membrane proteins to the mammalian endoplasmic reticulum is mediated by remarkably diverse signal sequences that have little or no homology with each other. Despite such sequence diversity, these signals are all recognized and interpreted by a highly conserved protein-conducting channel composed of the Sec61 complex. Signal recognition by Sec61 is essential for productive insertion of the nascent polypeptide into the translocation site, channel gating and initiation of transport. Although subtle differences in these steps can be detected between different substrates, it is not known whether they can be exploited to modulate protein translocation selectively. Here we describe cotransin, a small molecule that inhibits protein translocation into the endoplasmic reticulum. Cotransin acts in a signal-sequence-discriminatory manner to prevent the stable insertion of select nascent chains into the Sec61 translocation channel. Thus, the range of substrates accommodated by the channel can be specifically and reversibly modulated by a cell-permeable small molecule that alters the interaction between signal sequences and the Sec61 complex.

Amino Acid Sequence↗

Multiplicity and diversity of cloned zein cDNA sequences and their chromosomal localisation.

We have constructed and screened cDNA libraries from total maize endosperm poly(A) RNA or from a mRNA fraction enriched in zein sequences. From these libraries we have isolated clones representative of the major classes of zein cDNA sequences and have characterised them by crosshybridisation, by hybrid-selected translation, by in situ hybridisation to maize chromosomes, and hybridisation to genomic Southern blots. We conclude that at least four types of non cross-hybridising zein sequences are present, two coding for light chains and two for heavy chains. At least in the case of the light zeins, there is considerable sequence diversity among the clones which hybridise to each type. Similar results are obtained by translation of the mRNAs selected by each clone. In situ hybridisation shows that the light chain zein genes are located on chromosomes 4, 7, and 10, whilst genes coding for some of the heavy chain zeins are confined to the distal part of the long arm of chromosome 4.

Journal Article↗