Search PubMed⌕ Search

Biomedical subjects

Tanya Z Berardini

Publications and source records attributed to Tanya Z Berardini.

6 recordsLinked to original sources

PatMatch: a program for finding patterns in peptide and nucleotide sequences.

Here, we present PatMatch, an efficient, web-based pattern-matching program that enables searches for short nucleotide or peptide sequences such as cis-elements in nucleotide sequences or small domains and motifs in protein sequences. The program can be used to find matches to a user-specified sequence pattern that can be described using ambiguous sequence codes and a powerful and flexible pattern syntax based on regular expressions. A recent upgrade has improved performance and now supports both mismatches and wildcards in a single pattern. This enhancement has been achieved by replacing the previous searching algorithm, scan_for_matches [D'Souza et al. (1997), Trends in Genetics, 13, 497-498], with nondeterministic-reverse grep (NR-grep), a general pattern matching tool that allows for approximate string matching [Navarro (2001), Software Practice and Experience, 31, 1265-1312]. We have tailored NR-grep to be used for DNA and protein searches with PatMatch. The stand-alone version of the software can be adapted for use with any sequence dataset and is available for download at The Arabidopsis Information Resource (TAIR) at ftp://ftp.arabidopsis.org/home/tair/Software/Patmatch/. The PatMatch server is available on the web at http://www.arabidopsis.org/cgi-bin/patmatch/nph-patmatch.pl for searching Arabidopsis thaliana sequences.

Arabidopsis↗

Functional annotation of the Arabidopsis genome using controlled vocabularies.

Controlled vocabularies are increasingly used by databases to describe genes and gene products because they facilitate identification of similar genes within an organism or among different organisms. One of The Arabidopsis Information Resource's goals is to associate all Arabidopsis genes with terms developed by the Gene Ontology Consortium that describe the molecular function, biological process, and subcellular location of a gene product. We have also developed terms describing Arabidopsis anatomy and developmental stages and use these to annotate published gene expression data. As of March 2004, we used computational and manual annotation methods to make 85,666 annotations representing 26,624 unique loci. We focus on associating genes to controlled vocabulary terms based on experimental data from the literature and use The Arabidopsis Information Resource-developed PubSearch software to facilitate this process. Each annotation is tagged with a combination of evidence codes, evidence descriptions, and references that provide a robust means to assess data quality. Annotation of all Arabidopsis genes will allow quantitative comparisons between sets of genes derived from sources such as microarray experiments. The Arabidopsis annotation data will also facilitate annotation of newly sequenced plant genomes by using sequence similarity to transfer annotations to homologous genes. In addition, complete and up-to-date annotations will make unknown genes easy to identify and target for experimentation. Here, we describe the process of Arabidopsis functional annotation using a variety of data sources and illustrate several ways in which this information can be accessed and used to infer knowledge about Arabidopsis and other plant species.

Arabidopsis↗

Spatial clustering of isozyme-specific residues reveals unlikely determinants of isozyme specificity in fructose-1,6-bisphosphate aldolase.

Vertebrate fructose-1,6-bisphosphate aldolase exists as three isozymes (A, B, and C) that demonstrate kinetic properties that are consistent with their physiological role and tissue-specific expression. The isozymes demonstrate specific substrate cleavage efficiencies along with differences in the ability to interact with other proteins; however, it is unknown how these differences are conferred. An alignment of 21 known vertebrate aldolase sequences was used to identify all of the amino acids that are specific to each isozyme, or isozyme-specific residues (ISRs). The location of ISRs on the tertiary and quaternary structures of aldolase reveals that ISRs are found largely on the surface (24 out of 27) and are all outside of hydrogen bonding distance to any active site residue. Moreover, ISRs cluster into two patches on the surface of aldolase with one of these patches, the terminal surface patch, overlapping with the actin-binding site of aldolase A and overlapping an area of higher than average temperature factors derived from the x-ray crystal structures of the isozymes. The other patch, the distal surface patch, comprises an area with a different electrostatic surface potential when comparing isozymes. Despite their location distal to the active site, swapping ISRs between aldolase A and B by multiple site mutagenesis on recombinant expression plasmids is sufficient to convert the kinetic properties of aldolase A to those of aldolase B. This implies that ISRs influence catalysis via changes that alter the structure of the active site from a distance or via changes that alter the interaction of the mobile C-terminal portion with the active site. The methods used in the identification and analysis of ISRs discussed here can be applied to other protein families to reveal functionally relevant residue clusters not accessible by conventional primary sequence alignment methods.

Amino Acid Sequence↗

The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.

Arabidopsis thaliana is the most widely-studied plant today. The concerted efforts of over 11 000 researchers and 4000 organizations around the world are generating a rich diversity and quantity of information and materials. This information is made available through a comprehensive on-line resource called the Arabidopsis Information Resource (TAIR) (http://arabidopsis.org), which is accessible via commonly used web browsers and can be searched and downloaded in a number of ways. In the last two years, efforts have been focused on increasing data content and diversity, functionally annotating genes and gene products with controlled vocabularies, and improving data retrieval, analysis and visualization tools. New information include sequence polymorphisms including alleles, germplasms and phenotypes, Gene Ontology annotations, gene families, protein information, metabolic pathways, gene expression data from microarray experiments and seed and DNA stocks. New data visualization and analysis tools include SeqViewer, which interactively displays the genome from the whole chromosome down to 10 kb of nucleotide sequence and AraCyc, a metabolic pathway database and map tool that allows overlaying expression data onto the pathway diagrams. Finally, we have recently incorporated seed and DNA stock information from the Arabidopsis Biological Resource Center (ABRC) and implemented a shopping-cart style on-line ordering system.

Arabidopsis↗

HASTY, the Arabidopsis ortholog of exportin 5/MSN5, regulates phase change and morphogenesis.

Loss-of-function mutations of HASTY (HST) affect many different processes in Arabidopsis development. In addition to reducing the size of both roots and lateral organs of the shoot, hst mutations affect the size of the shoot apical meristem, accelerate vegetative phase change, delay floral induction under short days, adaxialize leaves and carpels, disrupt the phyllotaxis of the inflorescence, and reduce fertility. Double mutant analysis suggests that HST acts in parallel to SQUINT in the regulation of phase change and in parallel to KANADI in the regulation of leaf polarity. Positional cloning demonstrated that HST is the Arabidopsis ortholog of the importin beta-like nucleocytoplasmic transport receptors exportin 5 in mammals and MSN5 in yeast. Consistent with a potential role in nucleocytoplasmic transport, we found that HST interacts with RAN1 in a yeast two-hybrid assay and that a HST-GUS fusion protein is located at the periphery of the nucleus. HST is one of at least 17 members of the importin-beta family in Arabidopsis and is the first member of this family shown to have an essential function in plants. The hst loss-of-function phenotype suggests that this protein regulates the nucleocytoplasmic transport of molecules involved in several different morphogenetic pathways, as well as molecules generally required for root and shoot growth.

Amino Acid Sequence↗

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗