Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Improved database searches for orthologous sequences by conditioning on outgroup sequences.

MOTIVATION: Searches of biological sequence databases are usually focussed on distinguishing significant from random matches. However, the increasing abundance of related sequences on databases present a second challenge: to distinguish the evolutionarily most closely related sequences (often orthologues) from more distantly related homologues. This is particularly important when searching a database of partial sequences, where short orthologous sequences from a non-conserved region will score much more poorly than non-orthologous (outgroup) sequences from a conserved region. RESULTS: Such inferences are shown to be improved by conditioning the search results on the scores of an outgroup sequence. The log-odds score for each target sequence identified on the database has the log-odds score of the outgroup sequence subtracted from it. A test group of Caenorhabditis elegans kinase sequences and their identified C.elegans outgroups were searched against a test database of human Expressed Sequence Tag (EST) sequences, where the sets of true target sequences were known in advance. The outgroup conditioned method was shown to identify 58% more true positives ahead of the first false positive, compared to the straightforward search without an outgroup. A test dataset of 151 proteins drawn from the C.elegans genome, where the putative 'outgroup' was assigned automatically, similarly found 50% more true positives using outgroup conditioning. Thus, outgroup conditioning provides a means to improve the results of database searching with little increase in the search computation time.

Algorithms↗

Comparison of sequencing by hybridization and cycle sequencing for genotyping of human immunodeficiency virus type 1 reverse transcriptase.

The performances of two methods of nucleotide sequencing were compared for the detection of drug resistance mutations in human immunodeficiency virus type 1 reverse transcriptase (RT) in viruses isolated from highly RT inhibitor-experienced individuals. Of 11,677 amino acids deduced from population PCR products by both cycle sequencing and sequencing by hybridization to high-density arrays of oligonucleotide probes, 97.4% were concordant by both methods, 0.8% were discordant, and 1.7% had an ambiguous determination by at least one method. A higher rate of discordance (3.9%) was observed among RT inhibitor resistance-associated codons. In 45% of the isolates, RT codon 67 was deduced as the wild-type Asp by hybridization sequencing but as the zidovudine resistance-associated Asn by cycle sequencing. In other resistance-associated codon discordances, cycle sequencing also more commonly called a known resistance-associated amino acid than hybridization sequencing did. The nucleotide sequence in the vicinity of several codons with discordant calls influenced population-based hybridization sequencing. For isolates evaluated by additional sequencing of molecular clones of PCR products by both methods, the discordance between methods was less frequent (0.4% of all 5,994 amino acids and 0 of 494 drug resistance-associated codons). At positions which were discordant or ambiguous in the population sequences, the results of sequencing of clones by both methods were usually in agreement with the population cycle sequencing result. In summary, most RT codons were highly concordant by both methods of population-based sequencing, with discordances due in large part to genetic mixtures within or adjacent to discordant codons.

Anti-HIV Agents↗

AT-rich sequences containing Arabidopsis-type telomere sequence and their chromosomal distribution in Pinus densiflora.

Japanese red pine Pinus densiflora has 2 n=24 chromosomes and after FISH-detection of Arabidopsis-type (A-type) telomere sequences, many telomere signals were observed on these chromosomes at interstitial and proximal regions in addition to the chromosome ends. These interstitial and proximal signal sites were observed as DAPI-positive bands, suggesting that the interstitial and proximal telomere signal sites are composed of AT-rich highly repetitive sequences. Four DNA clones (PAL810, PAL1114, PAL1539, PAL1742) localized at the interstitial telomere signals were selected from AluI-digested genomic DNA library using colony blot hybridization probed with A-type telomere sequences and characterized using FISH and Southern blot hybridization. The AT-contents of these selected four clones were 60.8-76.3%, and repeat units of the telomere sequence and degenerated telomere sequences were found in their nucleotide sequences. Except for two sites of PAL1114, FISH signals of the four clones co-localized with interstitial and proximal A-type telomere sequence signals. FISH signals a showed similar distribution pattern, but the patterns of signal intensity were different among the four clones. PAL810, PAL1539 and PAL 1742 showed similar FISH signal patterns, and the differences were only with respect to the signal intensity of some signal sites. PAL1114 had unique signals that appeared on chromosomes 7 and 10. Based on results of the Southern blot hybridization these four sequences are not arranged tandemly. Our results suggest that the interstitial A-type telomere sequence signal sites were composed of a mixture of several AT-rich repetitive sequences and that these repetitive sequences contained A-type telomere sequences or degenerated A-type telomere sequence repeats.

AT Rich Sequence↗

ATP synthase from bovine mitochondria: complementary DNA sequence of the mitochondrial import precursor of the gamma-subunit and the genomic sequence of the mature protein.

The gamma-subunit of mitochondrial ATP synthase is part of the extrinsic membrane sector of the enzyme F1-ATPase. It is a nuclear gene product. Complementary DNA clones encoding a precursor of the protein have been isolated from a bovine library. The initial partial clone was identified with a mixture of 32 synthetic oligonucleotides designed from the known protein sequence (Walker et al., 1985), and this isolate was then used to screen the library again in order to find a complete cDNA. The DNA sequence of a clone that encodes the entire mature protein has been established, and the deduced protein sequence agrees exactly with that determined by direct sequence analysis of protein isolated from bovine hearts (Walker et al., 1985). At the 3' ends of two independently isolated clones, alternative polyadenylation sites have been observed; otherwise, the DNA sequences of the clones are concordant. In common with many other mitochondrial proteins encoded in nuclear genes, the deduced protein sequence has an N-terminal extension that is absent from the mature protein. These presequences direct the protein to its appropriate mitochondrial compartment and are removed during the import process. The cDNA clone has been employed to isolate bovine genomic clones containing the gene for the gamma-subunit. From them, the DNA sequence has been established of a region encoding the mature protein and six amino acids in the presequence, but not the remainder of the proposed import sequence. This sequence extends over almost 10 kb and is divided into eight exons. Intron B between exons I and II contains a sequence that is related to long interspersed repetitive elements (LINEs) that have been described in other mammals. Human LINEs are usually flanked by directly repeated sequences with a poly(A) tract at their 3' ends, and these features are present in the bovine LINE which is truncated. This sequence contains an open reading frame encoding part of a protein that is closely related to a protein encoded in mouse LINEs, to reverse transcriptase, and to DNA binding proteins. We have also made a preliminary investigation by DNA hybridization of the number of sequences related to the bovine gene in both the bovine and human genomes. Under the experimental conditions employed, one fragment hybridized in digests of bovine DNA, and two to four bands were detected in digests of human DNA; these latter fragments have originated from either expressed genes or pseudogenes.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Rapid and simple characterization of in vivo HIV-1 sequences using solid-phase direct sequencing.

Solid-phase direct sequencing was used to obtain in vivo sequence data of polymerase chain reaction (PCR)-amplified HIV-1 p25/p7 gene segments. The solid-phase sequencing method was compared to double-strand sequencing of preparative gel electrophoresis-purified amplification products and found to give more consistent results. Lysates of cells were compared to purified DNA as PCR template. HIV-1 sequences were as well amplified from lysates as from purified DNA and 7,780 bp of sequence from 41 samples were produced by direct sequencing. Sequence analysis revealed common sequence motifs relating the sequence to the Euroamerican and African groups of sequences previously described. The results indicate that many viruses of diverse origin circulate in Finland, although the majority seems to be of Euroamerican type. Solid-phase direct sequencing may provide a valuable tool for both epidemiological and pathogenic studies of in vivo HIV-1 infections.

Amino Acid Sequence↗

Nucleotide sequence and deduced amino acid sequence of the nonstructural proteins of dengue type 3 virus, Bangkok genotype.

The nucleotide sequence of the nonstructural protein gene (1,610 bases) of dengue 3 virus (Bangkok genotype; CH53489 isolated in 1973) has been determined in both forward and reverse directions. The PCR based cycle sequencing technic by the enzymatic method of Sanger et al using a sequencing primer 5'-end labeled with gamma-32P-ATP was the method of our choice for sequence analysis. Two cDNA templates were prepared by RT-PCR technique starting from the nucleotides 6,306-6,969 and 6,925-7,915 of the dengue 3 genome with the lengths of 663 and 990 base pairs respectively. In our cycle sequencing experiments, it has been observed that the substitution of 7-deaza-dG for dG in DNA eliminated most of the secondary structures that produce gel artifacts. The final sequence results of these two cDNA templates were established from their sequence data determined on both strands in opposite directions. Alignment between the newly established nucleotide sequences as well as their deduced amino acid sequences of the Bangkok dengue 3 (CH53489) virus and the published sequence data of the dengue 3 prototype (H87) was manipulated by the PC-DOS-GIBIO DNASIS TM 06-00 software. The homology of the nucleotide sequences between the two dengue 3 viruses was 96.65%. The deduced amino acid sequence from nucleotides 6,306-7,915 of the two viruses showed conserved amino acids of the nonstructural protein NS4a and 6 amino acid changes in NS4b and NS5.

Amino Acid Sequence↗

Identification of human chromosome 22 transcribed sequences with ORF expressed sequence tags.

Transcribed sequences in the human genome can be identified with confidence only by alignment with sequences derived from cDNAs synthesized from naturally occurring mRNAs. We constructed a set of 250,000 cDNAs that represent partial expressed gene sequences and that are biased toward the central coding regions of the resulting transcripts. They are termed ORF expressed sequence tags (ORESTES). The 250,000 ORESTES were assembled into 81,429 contigs. Of these, 1, 181 (1.45%) were found to match sequences in chromosome 22 with at least one ORESTES contig for 162 (65.6%) of the 247 known genes, for 67 (44.6%) of the 150 related genes, and for 45 of the 148 (30.4%) EST-predicted genes on this chromosome. Using a set of stringent criteria to validate our sequences, we identified a further 219 previously unannotated transcribed sequences on chromosome 22. Of these, 171 were in fact also defined by EST or full length cDNA sequences available in GenBank but not utilized in the initial annotation of the first human chromosome sequence. Thus despite representing less than 15% of all expressed human sequences in the public databases at the time of the present analysis, ORESTES sequences defined 48 transcribed sequences on chromosome 22 not defined by other sequences. All of the transcribed sequences defined by ORESTES coincided with DNA regions predicted as encoding exons by genscan. (http://genes.mit.edu/GENSCAN.html).

Chromosomes, Human, Pair 22↗

Fast assignment of protein structures to sequences using the intermediate sequence library PDB-ISL.

MOTIVATION: For large-scale structural assignment to sequences, as in computational structural genomics, a fast yet sensitive sequence search procedure is essential. A new approach using intermediate sequences was tested as a shortcut to iterative multiple sequence search methods such as PSI-BLAST. RESULTS: A library containing potential intermediate sequences for proteins of known structure (PDB-ISL) was constructed. The sequences in the library were collected from a large sequence database using the sequences of the domains of proteins of known structure as the query sequences and the program PSI-BLAST. Sequences of proteins of unknown structure can be matched to distantly related proteins of known structure by using pairwise sequence comparison methods to find homologues in PDB-ISL. Searches of PDB-ISL were calibrated, and the number of correct matches found at a given error rate was the same as that found by PSI-BLAST. The advantage of this library is that it uses pairwise sequence comparison methods, such as FASTA or BLAST2, and can, therefore, be searched easily and, in many cases, much more quickly than an iterative multiple sequence comparison method. The procedure is roughly 20 times faster than PSI-BLAST for small genomes and several hundred times for large genomes. AVAILABILITY: Sequences can be submitted to the PDB-ISL servers at http://stash.mrc-lmb.cam.ac.uk/PDB_ISL/ or http://cyrah.ebi.ac.uk:1111/Serv/PDB_ISL/ and can be downloaded from ftp://ftp.ebi.ac.uk/pub/contrib/jong/PDB_+ ++ISL/ CONTACT: sat@mrc-lmb.cam.ac.uk and jong@ebi.ac.uk

Peptide Library↗

Herpes simplex virus 1 reiterated S component sequences (c1) situated between the a sequence and alpha 4 gene are not essential for virus replication.

The herpes simplex virus 1 genome consists of two components, L and S, each containing unique sequences flanked by inverted repeats. Each of the 6.5-kilobase pair inverted repeats of the S component, designated a'c' and ca, contains an approximately 700-base pair sequence (designated c1) located between the a sequence and the 3' terminus of the alpha 4 gene. Like the a sequence, c1 consists of direct repeats and unique sequences. Its function is not known. To probe for its function, we constructed a plasmid containing a viral thymidine kinase (TK) gene inserted into the c1 sequence. The construct was recombined into the genome of a TK- virus by cotransfection with intact viral DNA and selection for TK+ virus. As predicted from previous studies (Knipe et al., Proc. Natl. Acad. Sci. U.S.A. 75:3896-3900, 1978), the TK gene was found to be present in both copies of the c1 sequence in the R3104 virus. To delete the c1 sequence we constructed a plasmid containing 4 kilobase pairs of pBR322 flanked by an a sequence and by structural sequences of the alpha 4 gene. In this instance the cells were transfected with the construct and R3104 DNA; the progeny of the transfection was plated in the presence of 5-bromo-2'-deoxyuridine, and the selection was for TK- virus (R3158). The pBR322 DNA sequences replaced the c1 at both termini of the S component in R3158 DNA, but a sequence homologous to c1 was present in proximity to the 3' terminus of the alpha 4 gene. The results indicate that the c1 region has no significant role in the replication of the virus in cell culture. The advantage of inserting the pBR322 sequence is that it permits efficient cloning of large herpes simplex virus 1 DNA fragments by simple ligation of digests and transformation of appropriate Escherichia coli strains. The effortless selection of recombinants carrying inserts in both copies of the c1 restates the usefulness of this technique for selection of insertion deletion recombinants and underscores the rapid emergence of sequence identity at both ends of the reiterated regions of the S component as previously reported (Knipe et al., Proc. Natl. Acad. Sci. U.S.A. 75:3896-3900, 1978).

DNA, Viral↗

Bridging expressed sequence alignments through targeted cDNA sequencing.

One of the major challenges in genome research is the identification of the complete set of genes in a genome. Alignments of expressed sequences (RNA and EST) with genomic sequences have been used to characterize genes. However, the number of alignments far exceeds the likely number of genes in a genome, suggesting that, for many genes, two or more alignments can be joined through overlapping sequences to yield accurate gene structures. High-throughput EST sequencing becomes less efficient in closing those alignment gaps due to its nonselective nature. We sought to bridge these alignments through a novel approach: targeted cDNA sequencing. Human expressed sequences from GenBank version 124 were aligned with the genomic sequence from NCBI build 24 using LEADS, Compugen's EST and RNA clustering and assembly software system. Nine hundred forty-eight pairs of alignments were selected based on EST clone information and/or their homology to the same known proteins. Reverse transcriptase PCR and sequencing yielded sequences for 363 of those pairs. These sequences helped characterize over 60 novel or otherwise incomplete genes in the recent UniGene build 153, which included over 1 million additional ESTs. These results indicate that this integrated and targeted strategy, combining computational prediction and experimental cDNA sequencing, can efficiently generate the overlapping sequences and enable the full characterization of genomes. Additional information about the contig pairs, the resultant overlapping sequences, tissue sources, and tissue profiles are available in a supplemental file.

Cloning, Molecular↗

Sequencing of HLA class I genes based on the conserved diversity of the noncoding regions: sequencing-based typing of the HLA-A gene.

We present a sequencing-based typing strategy for the HLA-A locus that is generally applicable to all HLA class I genes. Sequencing-based typing is the method of choice for matching in unrelated bone marrow transplantation on the allelic level. We determined the noncoding sequences of all serological antigens and most of their subtypes and discovered a remarkably conserved diversity characterized by polymorphic sequence motifs. In this study we took advantage of this diversity we uncovered in the 5' flanking region, 5' untranslated region and in the introns 1, 2 and 3, which was related to serological families. We established 12 primer mixes for setting up a PCR-based template preparation. Our strategy is based on the separate amplification of haplotypes and therefore defines the cis/trans linkage of polymorphic sequence motifs. This allowed individual sequencing of the haplotypes in all samples heterozygous for the broad antigens as well as the complete analysis of the polymorphic exons 2 and 3. All templates included the 2nd intron which was used as a priming site for the gene-specific 5' and 3' universal sequencing primers regardless of the amplified haplotypes. The independent sequencing of the haplotypes allows the application of the dye terminator cycle sequencing technique, which is less time-consuming and less-laborious than dye primer chemistry. The lack of heterozygous positions essentially facilitates on the one hand the data analysis and on the other hand the detection of new alleles. Sequencing is only required in one direction due to the absence of peak shift problems. The results will remain unambiguous regardless of a growing HLA sequence data bank since this sequencing technique defines the cis/trans linkage of sequence motifs in more than 95% of the cases.

Base Sequence↗

Genetic differences between blood- and brain-derived viral sequences from human immunodeficiency virus type 1-infected patients: evidence of conserved elements in the V3 region of the envelope protein of brain-derived sequences.

Human immunodeficiency virus type 1 (HIV-1) sequences were generated from blood and from brain tissue obtained by stereotactic biopsy from six patients undergoing a diagnostic neurosurgical procedure. Proviral DNA was directly amplified by nested PCR, and 8 to 36 clones from each sample were sequenced. Phylogenetic analysis of intrapatient envelope V3-V5 region HIV-1 DNA sequence sets revealed that brain viral sequences were clustered relative to the blood viral sequences, suggestive of tissue-specific compartmentalization of the virus in four of the six cases. In the other two cases, the blood and brain virus sequences were intermingled in the phylogenetic analyses, suggesting trafficking of virus between the two tissues. Slide-based PCR-driven in situ hybridization of two of the patients' brain biopsy samples confirmed our interpretation of the intrapatient phylogenetic analyses. Interpatient V3 region brain-derived sequence distances were significantly less than blood-derived sequence distances. Relative to the tip of the loop, the set of brain-derived viral sequences had a tendency towards negative or neutral charge compared with the set of blood-derived viral sequences. Entropy calculations were used as a measure of the variability at each position in alignments of blood and brain viral sequences. A relatively conserved set of positions were found, with a significantly lower entropy in the brain-than in the blood-derived viral sequences. These sites constitute a brain "signature pattern," or a noncontiguous set of amino acids in the V3 region conserved in viral sequences derived from brain tissue. This brain-derived signature pattern was also well preserved among isolates previously characterized in vitro as macrophage tropic. Macrophage-monocyte tropism may be the biological constraint that results in the conservation of the viral brain signature pattern.

Acquired Immunodeficiency Syndrome↗

The InDeVal insertion/deletion evaluation tool: a program for finding target regions in DNA sequences and for aiding in sequence comparison.

BACKGROUND: The program InDeVal was originally developed to help researchers find known regions of insertion/deletion activity (with the exception of isolated single-base indels) in newly determined Poaceae trnL-F sequences and compare them with 533 previously determined sequences. It is supplied with input files designed for this purpose. More broadly, the program is applicable for finding specific target regions (referred to as "variable regions") in DNA sequence. A variable region is any specific sequence fragment of interest, such as an indel region, a codon or codons, or sequence coding for a particular RNA secondary structure. RESULTS: InDeVal input is DNA sequence and a template file (sequence flanking each variable region). Additional files contain the variable regions and user-defined messages about the sequence found within them (e.g., taxa sharing each of the different indel patterns).Variable regions are found by determining the position of flanking sequence (referred to as "conserved regions") using the LPAM (Length-Preserving Alignment Method) algorithm. This algorithm was designed for InDeVal and is described here for the first time. InDeVal output is an interactive display of the analyzed sequence, broken into user-defined units. Once the user is satisfied with the organization of the display, the information can be exported to an annotated text file. CONCLUSIONS: InDeVal can find multiple variable regions simultaneously (28 indel regions in the Poaceae trnL-F files) and display user-selected messages specific to the sequence variants found. InDeVal output is designed to facilitate comparison between the analyzed sequence and previously evaluated sequence. The program's sensitivity to different levels of nucleotide and/or length variation in conserved regions can be adjusted. InDeVal is currently available for Windows in Additional file 1 or from http://www.sci.muni.cz/botany/elzdroje/indeval/.

Codon↗

Homologous sequences in steroidogenic enzymes, steroid receptors and a steroid binding protein suggest a consensus steroid-binding sequence.

The amino acid sequences of two steroidogenic enzymes, P450c17 (steroid 17 alpha-hydroxylase/17,20 lyse) and P450c21 (steroid 21-hydroxylase), are only 28.9% identical. However, these proteins share a region of 21 amino acids bearing 17 identical residues, which we previously suggested may represent the steroid binding site. We assembled a sequence database of known steroid-binding proteins and searched this with the sequence of this 21 amino acid region. The steroidogenic enzymes, P450c17, P450c21, P450scc (the cholesterol side-chain cleavage enzyme), and P450c11 (steroid 11 beta/18-hydroxylase) share a subregion of 17 amino acids having at least 15 identical residues. Related sequences were identified in a computerized search of the available sequences of steroid hormone receptors and binding proteins. These sequences were invariably found within larger domains previously associated with steroid binding. From these we propose a more general consensus sequence of LPLLL +/- 000KDRE0LKRL +/- PV, where +/- refers to any charged amino acid, and 0 refers to an uncharged amino acid. This consensus sequence predicts 147 or 187 total amino acids in 11 human proteins examined (78.6%). An equivalent degree of sequence identity, 178 of 221 amino acids (80.5%) was found among 13 animal homologs of these human proteins. The ability of this consensus sequence to predict 325 of 408 amino acids (79.7%) strongly suggests this sequence is necessary, if not sufficient, for a steroid binding site in many proteins. Lecithin-cholesterol acetyl transferase, cholesterol ester transfer protein, and steroid sulfatase did not have sequences similar to our consensus sequence.

Amino Acid Sequence↗

Studies of the hyperthermophile Thermotoga maritima by random sequencing of cDNA and genomic libraries. Identification and sequencing of the trpEG (D) operon.

Random sequencing of cDNA and genomic libraries has been used to study the genome of the hyperthermophile Thermotoga maritima. To date, 175 unique clones have been analyzed by comparing short sequence tags with known proteins in the PIR and GenBank databases. We find that a significant proportion of sequences can be matched to previously identified protein from non-Thermotoga sources. A high match rate was obtained from an oligo(dT)-primed cDNA library, where one-third of all unique sequences analyzed (21/65) shared high amino acid sequence similarity with proteins in the PIR and GenBank databases. Also, approximately one-third of the unique sequences from a second cDNA library (28/89), constructed with random oligo primers, could be matched to sequences in PIR and GenBank. Identification of genes from the oligo(dT)-primed cDNA library indicates that some Thermotoga mRNAs are polyadenylated. Genes have also been identified from a 1 to 2 kb genomic DNA library. Here, (3/21) of genomic sequences analyzed could be matched to protein in PIR and GenBank. One of the genomic clones had high sequence similarity to the tryptophan synthesis gene anthranilate synthase component I (trpE). Using this sequence tag, the Thermotoga trp operon was isolated and sequenced. The Thermotoga maritima trp operon is arranged with trpE forming an overlapping transcript with a second protein consisting of a fusion of anthranilate synthase component II (trpG) and anthranilate phosphoribosyltransferse (trpD). With regard to the fusion, the operon organization is similar to Escherichia coli and Salmonella typhimurium, but lacks the classic attenuation system of enteric bacteria. Amino acid sequence comparison with 19 trpE, 18 trpG and 14 trpD genes from other organisms suggest that the Thermotoga trp genes resemble corresponding genes from other thermophiles more closely than expected.

Amino Acid Sequence↗

Cloning and sequencing of the Bet v 1-homologous allergen Fra a 1 in strawberry (Fragaria ananassa) shows the presence of an intron and little variability in amino acid sequence.

The Fra a 1 allergen in strawberry (Fragaria ananassa) is homologous to the major birch pollen allergen Bet v 1, which has numerous isoforms differing in terms of amino acid sequence and immunological impact. To map the extent of sequence differences in the Fra a 1 allergen, PCR cloning and sequencing was applied. Several genomic sequences of Fra a 1, with a length of either 584, 591 or 594 nucleotides, were obtained from three different strawberry varieties. All contained one intron, with the length of either 101 or 110 nucleotides. By sequencing 30 different clones, eight different DNA sequences were obtained, giving in total five potential Fra a 1 protein isoforms, with high sequence similarity (>97% sequence identity) and only seven positions of amino acid variability, which were largely confirmed by mass spectrometry of expressed proteins. We conclude that the sequence variability in the strawberry allergen Fra a 1 is small, within and between strawberry varieties, and that multiple spots, previously detected in 2DE, are presumably due to differences in post-translational modification rather than differences in amino acid sequence. The most abundant Fra a 1 isoform sequence, recombinantly expressed in Escherichia coli after removal of the intron, was recognized by IgE from strawberry allergic patients. It cross-reacted with antibodies to Bet v 1 and the homologous apple allergen Mal d 1 (61 and 78% sequence identity, respectively), and will be used in further analyses of variation in Fra a 1-expression.

Allergens↗

Preliminary profile of the Cryptosporidium parvum genome: an expressed sequence tag and genome survey sequence analysis.

Cryptosporidium parvum is a protozoan enteropathogen that infects humans and animals and causes a pronounced diarrheal disease that can be life-threatening in immunocompromised hosts. No specific chemo- or immunotherapies exist to treat cryptosporidiosis and little molecular information is available to guide development of such therapies. To accelerate gene discovery and identify genes encoding potential drug and vaccine targets we constructed sporozoite cDNA and genomic DNA sequencing libraries from the Iowa isolate of C. parvum and determined approximately 2000 sequence tags by single-pass sequencing of random clones. Together, the 567 expressed sequence tags (ESTs) and 1507 genome survey sequences (GSSs) totaled one megabase (1 mb) of unique genomic sequence indicating that approximately 10% of the 10.4 mb C. parvum genome has been sequence tagged in this gene discovery expedition. The tags were used to search the public nucleic acid and protein databases via BLAST analyses, and 180 ESTs (32%) and 277 GSSs (18%) exhibited similarity with database sequences at smallest sum probabilities P(N)< or =10(-8). Some tags encoded proteins with clear therapeutic potential including S-adenosylhomocysteine hydrolase, histone deacetylase, polyketide/fatty-acid synthases, various cyclophilins, thrombospondin-related cysteine-rich protein and ATP-binding-cassette transporters. Several anonymous ESTs encoded proteins predicted to contain signal peptides or multiple transmembrane spanning segments suggesting they were destined for membrane-bound compartments, the cell surface or extracellular secretion. One-hundred four simple sequence repeats were identified within the nonredundant sequence tag collection with (TAA)(> or =6)/(TTA)(> or =6) and (TA)(> or = 10)/(AT)(> or =10 ) being the most prevalent, occurring 40 and 15 times, respectively. Various cellular RNAs and their genes were also identified including the small and large ribosomal RNAs, five tRNAs, the U2 small nuclear RNA, and the small and large virus-like, double-stranded RNAs. This investigation has demonstrated that survey sequencing is an efficient procedure for gene discovery and genome characterization and has identified and sequence tagged many C. parvum genes encoding potential therapeutic targets.

Amino Acid Sequence↗

Amplification of DNA sequences in wheat and its relatives: the Dgas44 and R350 families of repetitive sequences.

The sequence of a Triticum tauschii genomic clone representing a family of D-genome amplified DNA sequences, designated Dgas44, is reported. The Dgas44 sequence occurs on all chromosomes of the D genome of wheat, Triticum aestivum, and in situ hybridization revealed it to be evenly dispersed on all seven chromosome pairs. An internal HindIII fragment of Dgas44, designated Dgas44-3, defines the highly amplified region that is specific to the D genome. The polymerase chain reaction was used to amplify a 236-bp fragment within Dgas44-3 from chromosomes 1D, 2D, 3D, 4D, 5D, and 7D, and identical copies of this region of the Dgas44-3 sequence were found among the isolates from each of the chromosomes. The Dgas44-3 sequence population from specific chromosomes differed on average by 0.22% from the original Dgas44 sequence. The Dgas44 sequence was found to differentiate between the D genome present in T. aestivum, T. tauschii, hexaploid T. crassum, T. cylindricum, T. ventricosum, in which the sequence was present in a highly amplified form and T. juvenale, T. syriacum, and tetraploid T. crassum where the sequence family was difficult to detect. Another class of amplified sequences previously considered to be rye "specific." R350, was isolated from tetraploid wheat and its dispersed distribution on chromosomes was similar to the Dgas44 family in T. tauschii. In contrast with the Dgas44 sequence family, genome specificity for the remnant R350 sequence family was not evident since it was present on all wheat chromosomes.

Base Sequence↗