Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Excess nonsynonymous substitution of shared polymorphic sites among self-incompatibility alleles of Solanaceae.

The function of the self-incompatibility locus (S locus) of many plant species dictates that natural selection will favor high levels of protein diversity. Pairwise sequence comparisons between S alleles from four species of Solanaceae reveal remarkably high sequence diversity and evidence for shared polymorphism. The level of amino acid constraint was found to be significantly heterogeneous among different regions of the gene, with some regions being highly constrained and others appearing to be virtually unconstrained. In some regions of the protein, there was an excess of nonsynonymous over synonymous substitution, consistent with the strong diversifying selection that must operate on this locus. These hypervariable regions are candidates for the sites that determine functional allelic identity. Simple contingency table tests show that sites that have polymorphism shared between species have more nonsynonymous substitution than polymorphic sites that do not exhibit shared polymorphism. This is consistent with the idea that adaptive evolution favoring amino acid replacement is occurring at sites with shared polymorphism. Tests of clustered polymorphism reveal that an unusually low rate of recombination must be occurring in this locus, allowing very ancient alleles to preserve their identity.

Alleles↗

The prostate expression database (PEDB): status and enhancements in 2000.

The Prostate Expression Database (PEDB) is an online resource designed to access and analyze gene expression information derived from the human prostate. PEDB archives >55 000 expressed sequence tags (ESTs) from 43 cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library species comparisons. Recent enhancements to PEDB include: (i) the functional categorization of annotated EST assemblies using a classification scheme developed at The Institute for Genome Research; (ii) catalogs of expressed genes in specific prostate tissue sources designated as transcriptomes; and (iii) the addition of prostate proteome information derived from two-dimensional electrophoreses and mass spectrometry of prostate cancer cell lines. PEDB may be accessed via the WWW at http://www.mbt.washington.edu/PEDB/

Databases, Factual↗

Chimeric 16S rDNA sequences of diverse origin are accumulating in the public databases.

A significant number of chimeric 16S rDNA sequences of diverse origin were identified in the public databases by partial treeing analysis. This suggests that chimeric sequences, representing phylogenetically novel non-existent organisms, are routinely being overlooked in molecular phylogenetic surveys despite a general awareness of PCR-generated artefacts amongst researchers.

Chimera↗

mtDNA sequences suggest a recent evolutionary divergence for Beringian and northern North American populations.

Conventional descriptions of the pattern and process of human entry into the New World from Asia are incomplete and controversial. In order to gain an evolutionary insight into this process, we have sequenced the control region of mtDNA in samples of contemporary tribal populations of eastern Siberia, Alaska, and Greenland and have compared them with those of Amerind speakers of the Pacific Northwest and with those of the Altai of central Siberia. Specifically, we have analyzed sequence diversity in 33 mitochondrial lineages identified in 90 individuals belonging to five Circumpolar populations of Beringia, North America, and Greenland: Chukchi from Siberia, Inupiaq Eskimos and Athapaskans from Alaska, Eskimos from West Greenland, and Haida from Canada. Hereafter, we refer to these five populations as "Circumarctic peoples." These data were then compared with the sequence diversity in 47 mitochondrial lineages identified in a sample of 145 individuals from three Amerind-speaking tribes (Bella Coola, Nuu-Chah-Nulth, and Yakima) of the Pacific Northwest, plus 16 mitochondrial lineages identified in a sample of 17 Altai from central Siberia. Sequence diversity within and among Circumarctic populations is considerably less than the sequence diversity observed within and among the three Amerind tribes. The similarity of sequences found among the geographically dispersed Circumarctic groups, plus the small values of mean pairwise sequence differences within Circumarctic populations, suggest a recent and rapid evolutionary radiation of these populations. In addition, Circumarctic populations lack the 9-bp deletion which has been used to trace various migrations out of Asia, while populations of southeastern Siberia possess this deletion. On the basis of these observations, while the evolutionary affinities of Native Americans extend west to the Circumarctic populations of eastern Siberia, they do not include the Altai of central Siberia.

Alaska↗

Phylogenetic analysis of Theileria and Babesia equi in relation to the establishment of parasite populations within novel host species and the development of diagnostic tests.

The divergence of parasites is important for maintenance within an established host and spread to novel host species. In this paper we have carried out phylogenetic analyses of Theileria parasites isolated from different host species. This was performed with small subunit ribosomal RNA sequences available in the data bases and a novel sequence amplified from Theileria lestoquardi DNA. Similar phylogenetic studies were carried out with sequences representing the major merozoite/piroplasm surface antigen (mMPSA) from the data base, and novel sequences representing 2 mMPSA alleles from T. lestoquardi, a full length sequence of a Theileria taurotragi mMPSA gene and partial sequences of two new allelic variants of the Babesia equi mMPSA gene homologue. The analysis indicated that the pathogenic sheep parasite T. lestoquardi has most probably evolved from a common ancestor of T. annulata. Interestingly, the level of mMPSA sequence diversity found for T. lestoquardi was surprisingly low, while diversity between the B. equi sequences was higher than that found within any of the classical Theileria species. The possible implications of these results for the establishment of Theileria parasites within novel species are discussed. Extensive cross-reactivity of a range of antisera was found when tested against recombinant mMPSA polypeptides from different Theileria (including B. equi) species. The cross-reactivity between mMPSA polypeptides and sequence diversity are relevant for the development of species specific diagnostic tests.

Amino Acid Sequence↗

The creation of diversity in the human immunoglobulin V(lambda) repertoire.

Sequence diversity in the human antibody repertoire is generated in two steps: by the combinatorial assembly of V gene segments and by somatic hypermutation. Here, we have characterised these processes for the lambda (lambda) light chain using a library of 7600 lambda cDNA clones from peripheral blood lymphocytes. By hybridisation and sequencing we found that most lambda chains are derived from the cluster of V(lambda) segments closest to the J(lambda)-C(lambda) pairs and that there is considerable variation in the use of individual V(lambda) segments (ranging from 0.02% to 27%): three of the 30 functional V(lambda) segments encode half the expressed V(lambda) repertoire. As a result of these biases, sequence diversity in the primary repertoire is focused at the centre of the antigen binding site. By contrast, somatic hypermutation spreads diversity to the periphery. Comparison with the human kappa (kappa) light chain indicates that both kappa and lambda use the same strategy for searching sequence space and have almost identical patterns of diversity in the mature antibody repertoire.

Gene Frequency↗

Genetic variability of Crimean-Congo haemorrhagic fever virus in Russia and Central Asia.

Hyalomma marginatum ticks (449 pools, 4787 ticks in total) collected in European Russia and Dermacentor niveus ticks (100 pools, 1100 ticks in total) collected in Kazakhstan were screened by ELISA for the presence of Crimean-Congo haemorrhagic fever virus (CCHFV). Virus antigen was found in 10.2 and 3.0 % of the pools, respectively. RT-PCR was used to recover partial sequences of the CCHFV small (S) genome segment from seven pools of antigen-positive H. marginatum ticks, one pool of D. niveus ticks, four CCFH cases and four laboratory virus strains. Additionally, the entire S genome segments of the CCHFV strains STV/HU29223 (isolated from a patient in European Russia) and TI10145 (isolated from H. asiaticum in Uzbekistan) were amplified, cloned and sequenced. Phylogenetic analysis placed all CCHFV sequences from Russia in a single, well-supported clade (nucleotide sequence diversity up to 3.2 %). Virus sequences from H. marginatum were closely related or identical to those recovered from patients in the same regions of southern Russia. Newly described CCHFV strains from Central Asian countries fell into two genetic lineages. The first lineage was novel and included closely related virus sequences from Kazakhstan and Tajikistan (nucleotide sequence diversity up to 3.2 %). In contrast, a newly described CCHFV strain from Uzbekistan, strain TI10145, clustered on the phylogenetic trees with strains from China.

Animals↗

Assessment of hepatitis C virus sequence complexity by electrophoretic mobilities of both single-and double-stranded DNAs.

To assess genetic variation in hepatitis C virus (HCV) sequences accurately, we optimized a method for identifying distinct viral clones without determining the nucleotide sequence of each clone. Twelve serum samples were obtained from seven individuals soon after they acquired HCV during a prospective study, and a 452-bp fragment from the E2 region was amplified by reverse transcriptase PCR and cloned. Thirty-three cloned cDNAs representing each specimen were assessed by a method that combined heteroduplex analysis (HDA) and a single-stranded conformational polymorphism (SSCP) method to determine the number of clonotypes (electrophoretically indistinguishable cloned cDNAs) as a measure of genetic complexity (this combined method is referred to herein as the HDA+SSCP method). We calculated Shannon entropy, incorporating the number and distribution of clonotypes into a single quantifier of complexity. These measures were evaluated for their correlation with nucleotide sequence diversity. Blinded analysis revealed that the sensitivity (ability to detect variants) and specificity (avoidance of false detection) of the HDA+SSCP method were very high. The genetic distance (mean +/- standard deviation) between indistinguishable cloned cDNAs (intraclonotype diversity) was 0.6% +/- 0.9%, and 98.7% of cDNAs differed by <2%, while the mean distance between cloned cDNAs with different patterns was 4.0% +/- 3.2%. The sensitivity of the HDA+SSCP method compared favorably with either HDA or the SSCP method alone, which resulted in intraclonotype diversities of 1.6% +/- 1.8% and 3.5% +/- 3.4%, respectively. The number of clonotypes correlated strongly with genetic diversity (R2, 0.93), but this correlation fell off sharply when fewer clones were assessed. This HDA+SSCP method accurately reflected nucleotide sequence diversity among a large number of viral cDNA clones, which should enhance analyses to determine the effects of viral diversity on HCV-associated disease. If sequence diversity becomes recognized as an important parameter for staging or monitoring of HCV infection, this method should be practical enough for use in laboratories that perform nucleic acid testing.

Cloning, Molecular↗

Sequence and diversity of the rat delta T-cell receptor.

The cDNA sequence of the delta T-cell receptor (TCRD) in the adult Lewis rat thymus was determined using the technique of rapid amplification of cDNA ends. Sixteen variable region genes (TCRDV), two diversity regions (TCRDD), two joining regions (TCRDJ), and a single constant region gene (TCRDC) were identified. The sixteen unique TCRDV genes identified represented eight different subfamilies in the rat and were highly conserved (>80% nucleotide identity) to corresponding mouse sequences. Extensive junctional diversity was observed in the rat, with both TCRDD regions (TCRDD1 and TCRDD2) utilized in the majority of cDNA clones identified. The two TCRDJ genes were highly conserved and corresponded to TCRDJ1 and TCRDJ2 in the mouse; the majority of clones utilized TCRDJ1. The TCRDC region in the rat was 91.1% identical to the mouse TCRDC gene and was highly conserved to other species. Although extensive sequence information about mouse gamma-delta T-cell receptor genes is available, current knowledge of rat gamma-delta T-cells is limited. The sequence analysis presented in this study adds to our understanding of gamma-delta T-cells in general, and it may be utilized to study the role of gamma-delta T-cells in immune-mediated disease and transplantation models previously established in the rat.

Amino Acid Sequence↗

Human enteric Caliciviridae: a new prevalent small round-structured virus group defined by RNA-dependent RNA polymerase and capsid diversity.

Sequence comparison of the RNA-dependent RNA polymerases of small round-structured viruses (SRSVs) from 10 recent U.K. outbreaks of gastroenteritis revealed significant genetic variation. Computer analyses indicated that these viruses can be divided into two discrete groups. SRSV group I contains the previously characterized antigenic type 1 Norwalk and type 3 Southampton viruses. The amino acid sequences of the RNA polymerase, capsid and ORF3 of these two viruses are relatively similar (about 92%, 69% and 72% amino acid identity, respectively). A representative member of group II SRSVs, Bristol virus, was subjected to a detailed genetic analysis. Bristol virus is a recent antigenic type 2 isolate from a U.K. hospital outbreak of gastroenteritis. Using a single clinical sample the 3'-terminal 3881 nucleotide cDNA sequence [excluding the poly(A) tail] of this virus was determined. Analysis of the sequence revealed significant differences from those of group I viruses with the RNA polymerase region, capsid and ORF3 showing only about 62%, 43% and 30% amino acid identity respectively with the equivalent proteins of the Norwalk and Southampton viruses. These data suggest that the morphologically identical SRSVs belong to at least two genetically distinct groups.

Amino Acid Sequence↗

Culture-dependent and culture-independent diversity within the obligate marine actinomycete genus Salinispora.

Salinispora is the first obligate marine genus within the order Actinomycetales and a productive source of biologically active secondary metabolites. Despite a worldwide, tropical or subtropical distribution in marine sediments, only two Salinispora species have thus far been cultivated, suggesting limited species-level diversity. To further explore Salinispora diversity and distributions, the phylogenetic diversity of more than 350 strains isolated from sediments collected around the Bahamas was examined, including strains cultured using new enrichment methods. A culture-independent method, using a Salinispora-specific seminested PCR technique, was used to detect Salinispora from environmental DNA and estimate diversity. Overall, the 16S rRNA gene sequence diversity of cultured strains agreed well with that detected in the environmental clone libraries. Despite extensive effort, no new species level diversity was detected, and 97% of the 105 strains examined by restriction fragment length polymorphism belonged to one phylotype (S. arenicola). New intraspecific diversity was detected in the libraries, including an abundant new phylotype that has yet to be cultured, and a new depth record of 1,100 m was established for the genus. PCR-introduced error, primarily from Taq polymerase, significantly increased clone library sequence diversity and, if not masked from the analyses, would have led to an overestimation of total diversity. An environmental DNA extraction method specific for vegetative cells provided evidence for active actinomycete growth in marine sediments while indicating that a majority of sediment samples contained predominantly Salinispora spores at concentrations that could not be detected in environmental clone libraries. Challenges involved with the direct sequence-based detection of spore-forming microorganisms in environmental samples are discussed.

Bacteriological Techniques↗

The structural basis for the recognition of diverse receptor sequences by TRAF2.

Many members of the tumor necrosis factor receptor (TNFR) superfamily initiate intracellular signaling by recruiting TNFR-associated factors (TRAFs) through their cytoplasmic tails. TRAFs apparently recognize highly diverse receptor sequences. Crystal structures of the TRAF domain of human TRAF2 in complex with peptides from the TNFR family members CD40, CD30, Ox40, 4-1BB, and the EBV oncoprotein LMP1 revealed a conserved binding mode. A major TRAF2-binding consensus sequence, (P/S/A/T)x(Q/E)E, and a minor consensus motif, PxQxxD, can be defined from the structural analysis, which encompass all known TRAF2-binding sequences. The structural information provides a template for the further dissection of receptor binding specificity of TRAF2 and for the understanding of the complexity of TRAF-mediated signal transduction.

Amino Acid Sequence↗

Massive sequence perturbation of a small protein.

Most protein topologies rarely occur in nature, thus limiting our ability to extract sequence information that could be used to predict structure, function, and evolutionary constraints on protein folds. In principle, the sequence diversity explored by a given protein topology could be expanded by introducing sequence perturbations and selecting variant proteins that fold correctly. However, our capacity to explore sequence space is intrinsically limited by the enormous number of sequences generated from the 20 amino acids and the limited number of variants likely to fold. Here we sought to test whether the sequence space for naturally existing proteins can be explored by simple, sequential degeneration of a complete set of short sequence segments of a model protein, without long-range covariation. Using the Raf ras binding domain as a model of a small protein capable of autonomous folding, we degenerated 72 of 76 positions of the primary structure for the 20 amino acids in segments of four to seven residues defined by secondary structure and selected the folded species for interaction with h-ras by using an in vivo survival-selection assay. The methodology presented allowed for rigorous statistical analysis and comparison of sequence diversity. The ensemble of sequence variants of Raf ras binding domain obtained have recaptured the diversity observed for the ubiquitin-roll topology. A signature sequence for this fold and the implication of this strategy to protein design and structure prediction are discussed.

Amino Acid Sequence↗

Detection and identification of mammalian reoviruses in surface water by combined cell culture and reverse transcription-PCR.

Reoviruses are a common class of enteric viruses capable of infecting a broad range of mammalian species, typically with low pathogenicity. Previous studies have shown that reoviruses are common in raw water sources and are often found along with other animal viruses. This suggests that in addition to the commonly monitored enteroviruses, reoviruses might serve as an informative target for monitoring fecal contamination of drinking water sources. Mammalian reoviruses were detected and identified by a combined cell culture-reverse transcription-PCR (RT-PCR) assay with novel primers targeting the L3 gene that encodes the lambda3 major core protein. Five of 26 (19.2%) cytopathic effect-positive cell culture lysates inoculated with surface water were positive for reoviruses by RT-PCR. DNA sequence analysis of RT-PCR products revealed significant sequence diversity among isolates, which is consistent with the sequence diversity among previously characterized mammalian reoviruses. Sequence analysis revealed persistence of a reovirus genotype at a single sampling site, while a sample from another site contained two different reovirus genotypes.

Animals↗

Hypervariability within the Rifin, Stevor and Pfmc-2TM superfamilies in Plasmodium falciparum.

The human malaria parasite, Plasmodium falciparum, possesses a broad repertoire of proteins that are proposed to be trafficked to the erythrocyte cytoplasm or surface, based upon the presence within these proteins of a Pexel/VTS erythrocyte-trafficking motif. This catalog includes large families of predicted 2 transmembrane (2TM) proteins, including the Rifin, Stevor and Pfmc-2TM superfamilies, of which each possesses a region of extensive sequence diversity across paralogs and between isolates that is confined to a proposed surface-exposed loop on the infected erythrocyte. Here we express epitope-tagged versions of the 2TM proteins in transgenic NF54 parasites and present evidence that the Stevor and Pfmc-2TM families are exported to the erythrocyte membrane, thus supporting the hypothesis that host immune pressure drives antigenic diversity within the loop. An examination of multiple P.falciparum isolates demonstrates that the hypervariable loop within Stevor and Pfmc-2TM proteins possesses sequence diversity across isolate boundaries. The Pfmc-2TM genes are encoded within large amplified loci that share profound nucleotide identity, which in turn highlight the divergences observed within the hypervariable loop. The majority of Pexel/VTS proteins are organized together within sub-telomeric genome neighborhoods, and a mechanism must therefore exist to differentially generate sequence diversity within select genes, as well as within highly defined regions within these genes.

Amino Acid Sequence↗

Amino acid repeat patterns in protein sequences: their diversity and structural-functional implications.

All the protein sequences from SWISS-PROT database were analyzed for occurrence of single amino acid repeats, tandem oligo-peptide repeats, and periodically conserved amino acids. Single amino acid repeats of glutamine, serine, glutamic acid, glycine, and alanine seem to be tolerated to a considerable extent in many proteins. Tandem oligo-peptide repeats of different types with varying levels of conservation were detected in several proteins and found to be conspicuous, particularly in structural and cell surface proteins. It appears that repeated sequence patterns may be a mechanism that provides regular arrays of spatial and functional groups, useful for structural packing or for one to one interactions with target molecules. To facilitate further explorations, a database of Tandem Repeats in Protein Sequences (TRIPS) has been developed and is available at URL: http://www.ncl-india.org/trips.

Amino Acid Sequence↗

Comparison of Vpu sequences from diverse geographical isolates of HIV type 1 identifies the presence of highly variable domains, additional invariant amino acids, and a signature sequence motif common to subtype C isolates.

We compared the Vpu sequences from 101 strains of HIV-1 isolated from diverse geographical regions and various subtypes in order to identify regions of high variability, and those amino acid residues that were highly conserved or invariant. In addition to the highly conserved casein kinase II (CKII) phosphorylation sites, our analysis identified additional invariant residues in the transmembrane domain and in the first and second alpha-helical domains. Our analysis revealed that all subtype C sequences had a conserved LRLL motif at the C terminus that was also found in A/C intersubtype recombinants. While our analysis demonstrated the conservation of CKII domains in HIV-1 group M and O isolates, the number of potential CKII phosphorylation sites was variable in SIVcpz sequences. The results of this study will provide a basis for future mutagenesis studies to examine the role of certain amino acid residues in the structure and function of Vpu.

Amino Acid Motifs↗

Rapid evolution of two discrete regions of the caprine arthritis-encephalitis virus envelope surface glycoprotein during persistent infection.

Five major regions of sequence diversity between strains (V1-V5) have been described in the caprine arthritis-encephalitis lentivirus (CAEV) envelope surface unit glycoprotein (SU). To determine which of these variable regions is important in persistent infection in vivo, we evaluated SU sequence diversity in five neutralization variants from two goats and proviral DNA from five additional goats infected with CAEV-63 for up to 7 years. Overall amino acid sequence divergence in the SU encoded by provirus and neutralization variants compared to parental CAEV-63 ranged from 1.1 to 4%. However, most of the amino acid substitutions and all of the deletions and insertions were present in two discrete regions designated HV1 and HV2. The HV2 region was variable in all neutralization variants and provirus sequences from most animals. This region overlapped the V4 domain of CAEV SU and the neutralization domain of the closely related ovine maedi-visna lentivirus. HV1 was located in a region of SU strictly conserved in all small ruminant lentivirus strains except CAEV-63. This region only varied in a subset of neutralization variants and proviruses, all derived from goats with arthritis. In contrast, sequences in the V1,V2,V3, and V5 regions were stable in neutralization variants and proviruses from infected goats, indicating that sequence diversity between strains in these regions is not due to selection of variants in persistently infected animals. Our results define two discrete regions of CAEV SU that undergo rapid sequence variation in persistently infected goats which may have important roles in virus-host interactions.

Amino Acid Sequence↗