PubMed · 17088284
The TIGR Plant Transcript Assemblies database.
Abstract
The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kevin L Childs, John P Hamilton, Wei Zhu, Eugene Ly, Foo Cheung, Hank Wu, Pablo D Rabinowicz, Chris D Town, C Robin Buell, Agnes P Chan. 2006-11-06. The TIGR Plant Transcript Assemblies database.. https://doi.org/10.1093/nar%2Fgkl785
Cite the original work for its findings. Save a collection to share your selection of sources.