Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “third generation sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Structural and evolutionary relationships between two families of bacterial extracytoplasmic chaperone proteins which function cooperatively in fimbrial assembly.

Gram-negative purple bacteria possess pairs of extracytoplasmic, ATP-independent, fimbrium-specific chaperone proteins which cooperatively function in the assembly of this extracellular organelle. The two non-homologous families of these proteins have been termed "Fimbrial chaperone family no. 1" (FCF1) and "Fimbrial chaperone family no. 2" (FCF2). The eleven sequenced or partially sequenced members of each of these two protein families were analysed. Their sequences were multiply aligned, and average similarity and hydropathy plots were generated. Statistical analyses of the sequences revealed that the short FCF1 proteins (of about 240 residues) have been better conserved through evolutionary time than have the much larger FCF2 proteins (of about 830 residues). Moreover, the N-terminal thirds of the FCF2 proteins are better conserved than the central or C-terminal thirds of these proteins. Phylogenetic tree construction revealed that, in general, the two proteins which cooperate in the assembly of a particular fimbrial type have similar positions on their respective phylogenetic trees, suggesting that the two proteins evolved in parallel as a functional unit. Two exceptions were noted, however. In one case, a hybrid protein appears to have arisen, possibly by genetic recombination. In another case, the two proteins of a particular pair may have evolved separately and come together late in the evolutionary process to provide their cooperative function.

Amino Acid Sequence↗

recA-like genes from three archaean species with putative protein products similar to Rad51 and Dmc1 proteins of the yeast Saccharomyces cerevisiae.

The process of homologous recombination has been documented in bacterial and eucaryotic organisms. The Escherichia coli RecA and Saccharomyces cerevisiae Rad51 proteins are the archetypal members of two related families of proteins that play a central role in this process. Using the PCR process primed by degenerate oligonucleotides designed to encode regions of the proteins showing the greatest degree of identity, we examined DNA from three organisms of a third phylogenetically divergent group, Archaea, for sequences encoding proteins similar to RecA and Rad51. The archaeans examined were a hyperthermophilic acidophile, Sulfolobus sofataricus (Sso); a halophile, Haloferax volcanii (Hvo); and a hyperthermophilic piezophilic methanogen, Methanococcus jannaschii (Mja). The PCR generated DNA was used to clone a larger genomic DNA fragment containing an open reading frame (orf), that we refer to as the radA gene, for each of the three archaeans. As shown by amino acid sequence alignments, percent amino acid identities and phylogenetic analysis, the putative proteins encoded by all three are related to each other and to both the RecA and Rad51 families of proteins. The putative RadA proteins are more similar to the Rad51 family (approximately 40% identity at the amino acid level) than to the RecA family (approximately 20%). Conserved sequence motifs, putative tertiary structures and phylogenetic analysis implied by the alignment are discussed. The 5' ends of mRNA transcripts to the Sso radA were mapped. The levels of radA mRNA do not increase after treatment with UV irradiation as do recA and RAD51 transcripts in E.coli and S.cerevisiae. Hence it is likely that radA in this organism is a constitutively expressed gene and we discuss possible implications of the lack of UV-inducibility.

Amino Acid Sequence↗

Allelic drop-out may occur with a primer binding site polymorphism for the commonly used RFLP assay for the -1131T>C polymorphism of the Apolipoprotein AV gene.

Apolipoprotein AV (ApoAV) gene variant, -1131T>C, is associated with increased triglyceride concentrations in all ethnic groups studied. An MseI based RFLP analysis is the most commonly used method for genotyping this SNP. We genotyped a large cohort comprising 1185 Asian Indians and 173 UK Caucasians for -1131T>C using an ARMS-PCR based tetra-primer method. For quality control, we re-genotyped approximately 10% random samples from this cohort utilizing the MseI RFLP, which showed a 2.9% (3/102) genotyping error rate between the two methods. To investigate further, we sequenced the 900 bp region around the -1131T>C polymorphism in 25 Asian Indians and 15 UK Caucasians and found a number of polymorphisms including the -987C>T polymorphism. Further analysis of the -987C>T SNP showed a higher rare allele frequency of 0.23 in Asian Indians (n = 158) compared to 0.09 in the UK Caucasians (n = 157). This SNP is located 4 bp from the 3' end of the RFLP forward primer and is in weak linkage disequilibrium with -1131T>C variant (r2 = 0.084 and D' = 1). Repeated RFLP analysis of seven subjects heterozygous for -987C>T (seven times), showed discordant results with the sequence at -1131T>C SNP nearly one third (15/49) of the time. We conclude that presence of -987C>T polymorphism in the forward primer of the MseI RFLP assay may lead to allelic drop-out and generate unforeseen errors in genotyping the -1131T>C polymorphism. Our results also emphasise the need for careful quality control in all molecular genetic studies, particularly while transferring genotyping methods between various ethnic groups.

Alleles↗

Association of Fyn with the activated platelet-derived growth factor receptor: requirements for binding and phosphorylation.

Three members of the Src family of tyrosine kinases [pp60c-src (Src), p59fyn (Fyn) and pp62c-yes (Yes)] are ubiquitously expressed, and are thus likely to have general roles in growth control. We have previously shown that, after addition of platelet-derived growth factor (PDGF) to quiescent cells, all three kinases become activated and associated with the PDGF receptor. We have now addressed the requirements for this association. First, we have used a baculovirus expression system to show that Fyn associates with the activated PDGF receptor in vitro in the absence of other proteins, demonstrating that the association between the two molecules is direct. Second, by generating cell lines expressing chimeric molecules consisting of Fyn sequences fused to a portion of beta-galactosidase, we found that the SH2 domain of Fyn is necessary for ligand-stimulated association with the PDGF receptor in vivo. Third, those fusion proteins that associated with the PDGF receptor also became phosphorylated in vivo following PDGF treatment, and in in vitro kinase assays, suggesting that the amino-terminal half of Fyn contains the sites of PDGF-stimulated phosphorylation. Partially purified, kinase-negative Fyn also became phosphorylated in the activated PDGF receptor complex in vitro, demonstrating that the PDGF receptor phosphorylates Fyn, rather than the novel phosphorylations occurring by autophosphorylation.

Base Sequence↗

Solution-phase library screening for the identification of rare clones: isolation of an alpha 1D-adrenergic receptor cDNA.

alpha 1-Adrenergic receptor (alpha 1-AR) subtypes (alpha 1A and alpha 1B) play a critical role in vascular smooth muscle contraction and circulatory homeostasis. Transcripts for these guanine nucleotide-binding protein-coupled receptors are extremely low in abundance, however, and isolation of their cDNAs is difficult. We have developed a novel technique for identifying rare clones in a cDNA library, which has been used successfully to isolate a cDNA clone encoding an alpha 1D-AR. A 564-bp polymerase chain reaction product encoding a region between the third and sixth transmembrane domains of the alpha 1D-AR was first generated using rat brain mRNA as template and highly degenerate primers. The primers corresponded to those domains but contained mismatches to the alpha 1B-AR sequences. A 3-kb transcript was identified with this polymerase chain reaction probe, by Northern analysis of rat hippocampus. However, traditional plaque hybridization failed to identify a cDNA in a rat hippocampus lambda gt10 library. By solution-phase screening of virtually the entire library, a cDNA containing a 3-kb insert was identified, amplified, and purified. This insert encodes a 560-amino acid protein corresponding to the topology of guanine nucleotide-binding protein-coupled receptors. This receptor has approximately 71% amino acid identity, in the transmembrane regions, to the hamster and rat alpha 1B-ARs. Characterization of the receptor expressed in COS-7 cells, by ligand binding and photoaffinity labeling, revealed some of the characteristics of an alpha 1A-AR. However, unlike alpha 1A-ARs characterized previously in membrane preparations or in solubilized partially purified preparations, the expressed receptor could be extensively inactivated by chlorethylclonidine. In addition, it displays ligand-binding properties that are not consistent with an alpha 1A-AR. This indicates that the cDNA clone that we have isolated encodes a novel alpha 1-AR subtype, which we classify as the alpha 1D-AR.

Animals↗

A novel point mutation in the peripheral myelin protein 22 (PMP22) gene associated with Charcot-Marie-Tooth disease type 1A.

We studied the peripheral myelin protein gene PMP-22 in a large Sardinian family with Charcot-Marie-Tooth disease type 1A (CMT1A), in which the duplication commonly found in CMT1A was absent, but with evidence of linkage on chromosome 17. Sequencing of DNA and cDNA showed a missense point mutation G368-->T in exon 5 of PMP22, predicted to determine a valine for glycine substitution at codon 107, which could be plotted in the center of the PMP22 protein putative transmembrane domain III. Using sequence-specific oligonucleotide probes (SSOP), we found the point mutation in all affected CMT1A subjects but not in healthy family members or in 314 chromosomes of controls, thus indicating that the G368-->T point mutation is not a polymorphism. In the hypothetical model of PMP22, the amino acid at position 107 plots deeply into alpha-helical transmembrane domain III, a domain where point mutations have never previously been found. Although the same mutation was present in all CMT1A subjects examined, clinical findings showed a different stereotyped pattern in relation to the generation examined, for a progressive increase in severity and an earlier onset from the first to the third generation examined. Molecular analysis suggests that CMT1A disease in this family is due to the G368-->T point mutation, although other mechanisms may account for the clinical variability in the members of different generations.

Base Sequence↗

A novel alphaB-crystallin mutation associated with autosomal dominant congenital lamellar cataract.

PURPOSE: To identify the mutation and the underlying mechanism of cataractogenesis in a five-generation autosomal dominant congenital lamellar cataract family. METHODS: Nineteen mutation hot spots associated with autosomal dominant congenital cataract have been screened by PCR-based DNA sequencing. Recombinant wild-type and mutant human alphaB-crystallin were expressed in Escherichia coli and purified to homogeneity. The recombinant proteins were characterized by far UV circular dichroism, intrinsic tryptophan fluorescence, Bis-ANS fluorescence, multiangle light-scattering, and the measurement of chaperone activity. RESULTS: A novel missense mutation in the third exon of the alphaB-crystallin gene (CRYAB) was found to cosegregate with the disease phenotype in a five-generation autosomal dominant congenital lamellar cataract family. The single-base substitution (G-->A) results in the replacement of the aspartic acid residue by asparagine at codon 140. Far UV circular dichroism spectra indicated that the mutation did not significantly alter the secondary structure. However, intrinsic tryptophan fluorescence spectra and Bis-ANS fluorescence spectra indicated that the mutation resulted in alterations in tertiary and/or quaternary structures and surface hydrophobicity of alphaB-crystallin. Multiangle light-scattering measurement showed that the mutant alphaB-crystallin tended to aggregate into a larger complex than did the wild-type. The mutant alphaB-crystallin was more susceptible than wild-type to thermal denaturation. Furthermore, the mutant alphaB-crystallin not only lost its chaperone-like activity, it also behaved as a dominant negative which inhibited the chaperone-like activity of wild-type alphaB-crystallin. CONCLUSIONS: These data indicate that the altered tertiary and/or quaternary structures and the dominant negative effect of D140N mutant alphaB-crystallin underlie the molecular mechanism of cataractogenesis of this pedigree.

Adult↗

Variants of the Xenopus laevis ribosomal transcription factor xUBF are developmentally regulated by differential splicing.

XUBF is a Xenopus ribosomal transcription factor of the HMG-box family which contains five tandemly disposed homologies to the HMG1 & 2 DNA binding domains. XUBF has been isolated as a protein doublet and two cDNAs encoding the two molecular weight variants have been characterised. The major two forms of xUBF identified differ by the presence or absence of a 22 amino acid segment lying between HMG-boxes 3 and 4. Here we show that the mRNAs for these two forms of xUBF are regulated during development and differentiation over a range of nearly 20 fold. By isolating two of the xUBF genes, it was possible to show that both encoded the variable 22 amino acid segment in exon 12. Oocyte splicing assays and the sequencing of PCR-generated cDNA fragments, demonstrated that the transcripts from one of these genes were differentially spliced in a developmentally regulated manner. Transcripts from the second gene were found to be predominantly or exclusively spliced to produce the lower molecular weight form of xUBF. Expression of a high molecular weight form from yet a third gene was also detected. Although the intron-exon structures of the Xenopus and mouse UBF genes were found to be essentially identical, the differential splicing of exon 8 found in mammals, was not detected in Xenopus.

Amino Acid Sequence↗

An aminophospholipid translocase associated with body fat and type 2 diabetes phenotypes.

OBJECTIVE: We have shown that a region on proximal mouse chromosome 7, near the pink-eyed (p) dilution locus, contains an ATPase (pfatp), a putative aminophospholipid translocase. Studies have suggested that this gene is a prime candidate for modulating body fat or involved in lipid metabolism in mouse and humans. Toward further analyses, our objective was to generate the complete genomic structures of mouse and human genes. RESEARCH METHODS AND PROCEDURES: The genomic structure of mouse pfatp was deduced by comparing the full-length cDNA sequence with the genomic sequence derived from a mouse BAC. The human ortholog was identified from the National Center for Biotechnology Information database. Full-length cDNA was generated, and the corresponding genomic structure was deduced from the Human Genome Database. RESULTS: Murine pfatp, and its human ortholog, PFATP, belong to class V of the third subfamily of P-type ATPases. The gene organization is strikingly similar in both organisms and all exon-intron junctions are conserved. A putative promoter region of PFATP contains a strong CpG island. The 5' untranslated regions of the two cDNAs have potential binding sites for multiple transcription factors, including Sp1, USF, AP1, and AP2, involved in adipogenesis and adipocyte metabolism. DISCUSSION: We report the generation of the complete genomic structure of a novel aminophospholipid translocase in mice and humans. Because the exact biological role and the subsequent relevance of these ATPases to obesity and diabetes are unknown, these data help to delineate the role of these genes in lipid/adipocyte metabolism.

5' Untranslated Regions↗

An acidic cluster of human cytomegalovirus UL99 tegument protein is required for trafficking and function.

The human cytomegalovirus (HCMV) virion is comprised of a linear double-stranded DNA genome, proteinaceous capsid and tegument, and a lipid envelope containing virus-encoded glycoproteins. Of these components, the tegument is the least well defined in terms of both protein content and function. Several of the major tegument proteins are phosphoproteins (pp), including pp150, pp71, pp65, and pp28. pp28, encoded by the UL99 open reading frame (ORF), traffics to vacuole-like cytoplasmic structures and was shown recently to be essential for envelopment. To elucidate the UL99 amino acid sequences necessary for its trafficking and function in the HCMV replication cycle, two types of viral mutants were analyzed. Using a series of recombinant viruses expressing various UL99-green fluorescent protein fusions, we demonstrate that myristoylation at glycine 2 and an acidic cluster (AC; amino acids 44 to 57) are required for the punctate perinuclear and cytoplasmic (vacuole-like) localization observed for wild-type pp28. A second approach involving the generation of several UL99 deletion mutants indicated that at least the C-terminal two-thirds of this ORF is nonessential for viral growth. Furthermore, the data suggest that an N-terminal region of UL99 containing the AC is required for viral growth. Regarding virion incorporation or UL99-encoded proteins, we provide evidence that suggests that a hypophosphorylated form of pp28 is incorporated, myristoylation is required, and sequences within the first 57 amino acids are sufficient.

Amino Acid Sequence↗

Saposin A: second cerebrosidase activator protein.

Saposin A, a heat-stable 16-kDa glycoprotein, was isolated from Gaucher disease spleen and purified to homogeneity. Chemical sequencing from its amino terminus and of peptides obtained by digestion with protease from Staphylococcus aureus strain V-8 demonstrated that saposin A is derived from proteolytic processing of domain 1 of its precursor protein, prosaposin. Processing of prosaposin (70 kDa) also generates three other previously reported saposin proteins, B, C, and D, from its second, third, and fourth domains. Similar to saposin C, saposin A stimulates the hydrolysis of 4-methylumbelliferyl beta-glucoside and glucocerebroside by beta-glucosylceramidase and of galactocerebroside by beta-galactosylceramidase, mainly by increasing the maximal velocity of both reactions. Saposin A is as active as saposin C in these reactions. Saposin A has no significant effect on other sphingolipid and 4-methylumbelliferyl glycoside hydrolases tested. Saposin A has two potential glycosylation sites that appear to be glycosylated. After deglycosylation, saposin A had a subunit molecular mass of 10 kDa and was as active as native saposin A. However, reduction and alkylation abolished the activation. A three-dimensional model comparing saposins A and C reveals significant sequence homology between them, especially preservation of conserved acidic and basic residues in their middle regions. Each appears to possess a conformationally rigid hydrophobic pocket stabilized by three internal disulfide bridges, with amphipathic helical regions interrupted by helix breakers.

Amino Acid Sequence↗

The gene cluster inlC2DE of Listeria monocytogenes contains additional new internalin genes and is important for virulence in mice.

In this work we identified and characterized a gene cluster containing three internalin genes of Listeria monocytogenes EGD. These genes, termed inlG, inlH and inlE, encode proteins of 490, 548 and 499 amino acids, respectively, which belong to the family of large, cell wall-bound internalins. The inlGHE gene cluster is flanked by two listerial house-keeping genes encoding proteins homologous to the 6-phospho-beta-glucosidase and the succinyl-diaminopimelate desuccinylase of E. coli. A similar internalin gene cluster, inlC2DE, localised to the same position on the L. monocytogenes EGD chromosome was recently described in a different isolate (Dramsi S, Dehoux P, Lebrun M, Goossens PL, Cossart P (1997) Infect Immun 65: 1615-1625). Sequence comparison of the two inl gene clusters indicates that inlG is a new internalin gene, while inlH was generated by a site-specific recombination, leading to an in-frame deletion which removed the 3'-terminal end of inlC2 and the 5'-terminal part of inlD. The third gene of the inlGHE cluster, inlE, is almost identical to the previously reported inlE gene. Our data show that the inlGHE gene cluster is probably transcribed from a major PrfA-independent promoter located upstream of inlG. PCR analysis revealed the presence of the newly identified inl genes inlG and inlH in most L. monocytogenes isolates tested. A mutant which has lost inlG, inlH and inlE by an in-frame deletion exhibited, after oral infection of mice, a significant loss in virulence and shows drastically reduced numbers of viable bacteria in both liver and spleen when compared to the wild-type strain.

Amino Acid Sequence↗

Antigenic subclasses of polytropic murine leukemia virus (MLV) isolates reflect three distinct groups of endogenous polytropic MLV-related sequences in NFS/N mice.

Polytropic murine leukemia viruses (MLVs) are generated by recombination of ecotropic MLVs with members of a family of endogenous proviruses in mice. Previous studies have indicated that polytropic MLV isolates comprise two mutually exclusive antigenic subclasses, each of which is reactive with one of two monoclonal antibodies termed MAb 516 and Hy 7. A major determinant of the epitopes distinguishing the subclasses mapped to a single amino acid difference in the SU protein. Furthermore, distinctly different populations of the polytropic MLV subclasses are generated upon inoculation of different ecotropic MLVs. Here we have characterized the majority of endogenous polytropic MLV-related proviruses of NFS/N mice. Most of the proviruses contain intact sequences encoding the receptor-binding region of the SU protein and could be distinguished by sequence heterogeneity within that region. We found that the endogenous proviruses comprise two major groups that encode the major determinant for Hy 7 or MAb 516 reactivity. The Hy 7-reactive proviruses correspond to previously identified polytropic proviruses, while the 516-reactive proviruses comprise the modified polytropic proviruses as well as a third group of polytropic MLV-related proviruses that exhibit distinct structural features. Phylogenetic analyses indicate that the latter proviruses reflect features of phylogenetic intermediates linking xenotropic MLVs to the polytropic and modified polytropic proviruses. These studies elucidate the relationships of the antigenic subclasses of polytropic MLVs to their endogenous counterparts, identify a new group of endogenous proviruses, and identify distinguishing characteristics of the proviruses that should facilitate a more precise description of their expression in mice and their participation in recombination to generate recombinant viruses.

Animals↗

Regulation, function and potential origin of the Drosophila gene spalt adjacent, which encodes a secreted protein expressed in the early embryo.

During early embryogenesis of Drosophila the spatial and temporal expression patterns of the region-specific homeotic gene spalt (sal) and the neighbouring gene spalt adjacent (sala) extensively overlap. We show that the initial expression patterns of the two genes in the blastoderm also have identical genetic controls. However, while sal encodes a transcription factor, sala encodes a precursor protein from which a functional signal peptide is cleaved off to generate the secreted sala protein. Ectopic expression or absence of sala protein does not affect embryonic development, adult viability or fertility. In addition to sal and sala, we identified a third gene nearby, termed spalt related (salr), which shares coding sequence similarity and a late embryonic expression pattern with sal, but lacks the early expression domains that are shared by sal and sala. These results suggest that the three genes and their present cis-regulatory regions arose through a chromosomal rearrangement involving local duplication and transposition events in the 32F/33A region on the left arm of the second chromosome.

Amino Acid Sequence↗

Identification and in vivo expression of a prokaryotic-like ribosome recognition sequence upstream of the coat protein gene of potato virus X.

Conserved prokaryotic sequence motifs, distinct from the classic Shine-Dalgarno sequence, yet possessing homology to 16S rRNA in E. coli have been identified in a number of plant viruses. In this report, a similar Shine-Dalgarno-like motif located immediately upstream to the CP gene of potato virus X (PVX) was demonstrated to enable expression of a reporter gene in E. coli to approximately one third the level of a similar construct containing the classical Shine-Dalgarno sequence. Both PVX-specific CP and RNA transcripts were detected in chloroplasts purified from transgenic potato plants containing the PVX CP gene and corresponding leader sequence. Protoplasts generated from these transgenic plants were used to demonstrate that expression of the PVX CP from chloroplasts is possible. The implications of these results on the PVX infection cycle are discussed.

Base Sequence↗

Branched DNA signal amplification for direct quantitation of nucleic acid sequences in clinical specimens.

In this chapter I have reviewed the development of bDNA as a method for quantitation of nucleic acid targets and the application of this technology to the study of infectious diseases and cell biology. The ability to quantify viral nucleic acids in clinical specimens has led to a better understanding of the pathogenesis of chronic viral infections such as HIV-1, HCV, and HBV. The information provided by these methods can also be important in the management of patients with these infections. The prognostic value of a single baseline HIV-1 RNA level rivals that surgical staging procedures for cancer, which are among the most powerfully predictive tests in medicine (Mellors et al., 1996). These methods have been used to assess rapidly the effects of antiviral therapy, which has both expedited the development of antiviral drugs and improved the management of patients with HIV-1 and HCV infections. bDNA has several characteristics that distinguish it from the quantitative target amplification systems, including better tolerance of target sequence variability, more direct measurement of target, simpler sample preparation, and less sample-to-sample variation. However, the first- and second-generation bDNA assays lacked sensitivity compared with the target amplifications systems. The changes incorporated into the third-generation assays have effectively increased the signal-to-noise ratio to such a high level that the analytical sensitivity of system 8 bDNA approaches that of PCR. In theory, bDNA can be made even more sensitive by increasing both the sample volume and the signal-to-noise ratio. Nonspecific hybridization can be further reduced by finding more effective blockers for the solid phase or by redesigning the amplifier molecule or the solid phase itself. The increased sensitivity may create new applications for the technology in filter and in situ hybridization assays.

Animals↗

Telomere elongation (Tel), a new mutation in Drosophila melanogaster that produces long telomeres.

In most eukaryotes telomeres are extended by telomerase. Drosophila melanogaster, however, lacks telomerase, and telomere-specific non-LTR retrotransposons, HeT-A and TART, transpose specifically to chromosome ends. A Drosophila strain, Gaiano, that has long telomeres has been identified. We extracted the major Gaiano chromosomes into an Oregon-R genetic background and examined the resulting stocks after 60 generations. In situ hybridization using HeT-A and TART sequences showed that, in stocks carrying either the X or the second chromosome from Gaiano, only the Gaiano-derived chromosomes display long telomeres. However, in stocks carrying the Gaiano third chromosome, all telomeres are substantially elongated, indicating that the Gaiano chromosome 3 carries a factor that increases HeT-A and TART addition to the telomeres. We show that this factor, termed Telomere elongation (Tel), is dominant and localizes as a single unit to 69 on the genetic map. The long telomeres tend to associate with each other in both polytene and mitotic cells. These associations depend on telomere length rather than the presence of Tel. Associations between metaphase chromosomes are resolved during anaphase, suggesting that they are mediated by either proteinaceous links or DNA hydrogen bonding, rather than covalent DNA-DNA bonds.

Anaphase↗

CR-EST: a resource for crop ESTs.

The crop expressed sequence tag database, CR-EST (http://pgrc.ipk-gatersleben.de/cr-est/), is a publicly available online resource providing access to sequence, classification, clustering and annotation data of crop EST projects. CR-EST currently holds more than 200,000 sequences derived from 41 cDNA libraries of four species: barley, wheat, pea and potato. The barley section comprises approximately one-third of all publicly available ESTs. CR-EST deploys an automatic EST preparation pipeline that includes the identification of chimeric clones in order to transparently display the data quality. Sequences are clustered in species-specific projects to currently generate a non-redundant set of approximately 22,600 consensus sequences and approximately 17,200 singletons, which form the basis of the provided set of unigenes. A web application allows the user to compute BLAST alignments of query sequences against the CR-EST database, query data from Gene Ontology and metabolic pathway annotations and query sequence similarities from stored BLAST results. CR-EST also features interactive JAVA-based tools, allowing the visualization of open reading frames and the explorative analysis of Gene Ontology mappings applied to ESTs.

Crops, Agricultural↗