Search PubMed⌕ Search

Biomedical subjects

Steven E Brenner

Publications and source records attributed to Steven E Brenner.

35 records · Page 2Linked to original sources

Genome-wide analysis reveals an unexpected function for the Drosophila splicing factor U2AF50 in the nuclear export of intronless mRNAs.

The protein factor U2AF is an essential component required for pre-mRNA splicing. Mutations identified in the S. pombe large U2AF subunit were used to engineer transgenic Drosophila carrying temperature-sensitive U2AF large subunit alleles. Mutant recombinant U2AF heterodimers showed reduced polypyrimidine tract RNA binding at elevated temperatures. Genome-wide RNA profiling comparing wild-type and mutant strains identified more than 400 genes differentially expressed in the dU2AF50 mutant flies grown at the restrictive temperature. Surprisingly, almost 40% of the downregulated genes lack introns. Microarray analyses revealed that nuclear export of a large number of intronless mRNAs is impaired in Drosophila-cultured cells RNAi knocked down for dU2AF50. Immunopurification of nuclear RNP complexes showed that dU2AF50 associates with intronless mRNAs. These results reveal an unexpected role for the splicing factor dU2AF50 in the nuclear export of intronless mRNAs.

Active Transport, Cell Nucleus↗

Structural studies of the Nudix hydrolase DR1025 from Deinococcus radiodurans and its ligand complexes.

We have determined the crystal structure, at 1.4A, of the Nudix hydrolase DR1025 from the extremely radiation resistant bacterium Deinococcus radiodurans. The protein forms an intertwined homodimer by exchanging N-terminal segments between chains. We have identified additional conserved elements of the Nudix fold, including the metal-binding motif, a kinked beta-strand characterized by a proline two positions upstream of the Nudix consensus sequence, and participation of the N-terminal extension in the formation of the substrate-binding pocket. Crystal structures were also solved of DR1025 crystallized in the presence of magnesium and either a GTP analog or Ap(4)A (both at 1.6A resolution). In the Ap(4)A co-crystal, the electron density indicated that the product of asymmetric hydrolysis, ATP, was bound to the enzyme. The GTP analog bound structure showed that GTP was bound almost identically as ATP. Neither nucleoside triphosphate was further cleaved.

Adenosine Triphosphate↗

Three-dimensional motifs from the SCOR, structural classification of RNA database: extruded strands, base triples, tetraloops and U-turns.

Release 2.0.1 of the Structural Classification of RNA (SCOR) database, http://scor.lbl.gov, contains a classification of the internal and hairpin loops in a comprehensive collection of 497 NMR and X-ray RNA structures. This report discusses findings of the classification that have not been reported previously. The SCOR database contains multiple examples of a newly described RNA motif, the extruded helical single strand. Internal loop base triples are classified in SCOR according to their three-dimensional context. These internal loop triples contain several examples of a frequently found motif, the minor groove AGC triple. SCOR also presents the predominant and alternate conformations of hairpin loops, as shown in the most well represented tetraloops, with consensus sequences GNRA, UNCG and ANYA. The ubiquity of the GNRA hairpin turn motif is illustrated by its presence in complex internal loops.

Base Pairing↗

Protein secondary structure: entropy, correlations and prediction.

MOTIVATION: Is protein secondary structure primarily determined by local interactions between residues closely spaced along the amino acid backbone or by non-local tertiary interactions? To answer this question, we measure the entropy densities of primary and secondary structure sequences, and the local inter-sequence mutual information density. RESULTS: We find that the important inter-sequence interactions are short ranged, that correlations between neighboring amino acids are essentially uninformative and that only one-fourth of the total information needed to determine the secondary structure is available from local inter-sequence correlations. These observations support the view that the majority of most proteins fold via a cooperative process where secondary and tertiary structure form concurrently. Moreover, existing single-sequence secondary structure prediction algorithms are almost optimal, and we should not expect a dramatic improvement in prediction accuracy. AVAILABILITY: Both the data sets and analysis code are freely available from our Web site at http://compbio.berkeley.edu/

Algorithms↗

An unappreciated role for RNA surveillance.

BACKGROUND: Nonsense-mediated mRNA decay (NMD) is a eukaryotic mRNA surveillance mechanism that detects and degrades mRNAs with premature termination codons (PTC+ mRNAs). In mammals, a termination codon is recognized as premature if it lies more than about 50 nucleotides upstream of the final intron position. More than a third of reliably inferred alternative splicing events in humans have been shown to result in PTC+ mRNA isoforms. As the mechanistic details of NMD have only recently been elucidated, we hypothesized that many PTC+ isoforms may have been cloned, characterized and deposited in the public databases, even though they would be targeted for degradation in vivo. RESULTS: We analyzed the human alternative protein isoforms described in the SWISS-PROT database and found that 144 (5.8% of 2,483) isoform sequences amenable to analysis, from 107 (7.9% of 1,363) SWISS-PROT entries, derive from PTC+ mRNA. CONCLUSIONS: For several of the PTC+ isoforms we identified, existing experimental evidence can be reinterpreted and is consistent with the action of NMD to degrade the transcripts. Several genes with mRNA isoforms that we identified as PTC+--calpain-10, the CDC-like kinases (CLKs) and LARD--show how previous experimental results may be understood in light of NMD.

Alternative Splicing↗

The ASTRAL Compendium in 2004.

The ASTRAL Compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. Partially derived from the SCOP database of protein structure domains, it includes sequences for each domain and other resources useful for studying these sequences and domain structures. The current release of ASTRAL contains 54,745 domains, more than three times as many as the initial release 4 years ago. ASTRAL has undergone major transformations in the past 2 years. In addition to several complete updates each year, ASTRAL is now updated on a weekly basis with preliminary classifications of domains from newly released PDB structures. These classifications are available as a stand-alone database, as well as integrated into other ASTRAL databases such as representative subsets. To enhance the utility of ASTRAL to structural biologists, all SCOP domains are now made available as PDB-style coordinate files as well as sequences. In addition to sequences and representative subsets based on SCOP domains, sequences and subsets based on PDB chains are newly included in ASTRAL. Several search tools have been added to ASTRAL to facilitate retrieval of data by individual users and automated methods. ASTRAL may be accessed at http://astral.stanford. edu/.

Animals↗

SCOP database in 2004: refinements integrate structure and sequence family data.

The Structural Classification of Proteins (SCOP) database is a comprehensive ordering of all proteins of known structure, according to their evolutionary and structural relationships. Protein domains in SCOP are hierarchically classified into families, superfamilies, folds and classes. The continual accumulation of sequence and structural data allows more rigorous analysis and provides important information for understanding the protein world and its evolutionary repertoire. SCOP participates in a project that aims to rationalize and integrate the data on proteins held in several sequence and structure databases. As part of this project, starting with release 1.63, we have initiated a refinement of the SCOP classification, which introduces a number of changes mostly at the levels below superfamily. The pending SCOP reclassification will be carried out gradually through a number of future releases. In addition to the expanded set of static links to external resources, available at the level of domain entries, we have started modernization of the interface capabilities of SCOP allowing more dynamic links with other databases. SCOP can be accessed at http://scop.mrc-lmb.cam.ac.uk/scop.

Animals↗

SCOR: Structural Classification of RNA, version 2.0.

SCOR, the Structural Classification of RNA (http://scor.lbl.gov), is a database designed to provide a comprehensive perspective and understanding of RNA motif three-dimensional structure, function, tertiary interactions and their relationships. SCOR 2.0 represents a major expansion and introduces a new classification organization. The new version represents the classification as a Directed Acyclic Graph (DAG), which allows a classification node to have multiple parents, in contrast to the strictly hierarchical classification used in SCOR 1.2. SCOR 2.0 supports three types of query terms in the updated search engine: PDB or NDB identifier, nucleotide sequence and keyword. We also provide parseable XML files for all information. This new release contains 511 RNA entries from the PDB as of 15 May 2003. A total of 5880 secondary structural elements are classified: 2104 hairpin loops and 3776 internal loops. RNA motifs reported in the literature, such as 'Kink turn' and 'GNRA loops', are now incorporated into the structural classification along with definitions and descriptions.

Animals↗

The evolving roles of alternative splicing.

Alternative splicing is now commonly thought to affect more than half of all human genes. Recent studies have investigated not only the scope but also the biological impact of alternative splicing on a large scale, revealing that its role in generating proteome diversity may be augmented by a role in regulation. For instance, protein function can be regulated by the removal of interaction or localization domains by alternative splicing. Alternative splicing can also regulate gene expression by splicing transcripts into unproductive mRNAs targeted for degradation. To fully understand the scope of alternative splicing, we must also determine how many of the predicted splice variants represent functional forms. Comparisons of alternative splicing between human and mouse genes show that predominant splice variants are usually conserved, but rare variants are less commonly shared. Evolutionary conservation of splicing patterns suggests functional importance and provides insight into the evolutionary history of alternative splicing.

Alternative Splicing↗

WebLogo: a sequence logo generator.

WebLogo generates sequence logos, graphical representations of the patterns within a multiple sequence alignment. Sequence logos provide a richer and more precise description of sequence similarity than consensus sequences and can rapidly reveal significant features of the alignment otherwise difficult to perceive. Each logo consists of stacks of letters, one stack for each position in the sequence. The overall height of each stack indicates the sequence conservation at that position (measured in bits), whereas the height of symbols within the stack reflects the relative frequency of the corresponding amino or nucleic acid at that position. WebLogo has been enhanced recently with additional features and options, to provide a convenient and highly configurable sequence logo generator. A command line interface and the complete, open WebLogo source code are available for local installation and customization.

Amino Acid Sequence↗

Widespread predicted nonsense-mediated mRNA decay of alternatively-spliced transcripts of human normal and disease genes.

We have recently shown that a third of reliably-inferred alternative mRNA isoforms are candidates for nonsense-mediated mRNA decay (NMD), an mRNA surveillance system (Lewis et al., 2003; PROC: Natl Acad. Sci. USA, 100, 189-192). Rather than being translated to yield protein, these transcripts are expected to be degraded and may be subject to regulated unproductive splicing and translation (RUST). Our initial experimental studies are consistent with these predictions and suggest an unappreciated role for NMD in several human diseases.

Alternative Splicing↗

Evidence for the widespread coupling of alternative splicing and nonsense-mediated mRNA decay in humans.

To better understand the role of alternative splicing, we conducted a large-scale analysis of reliable alternative isoforms of known human genes. Each isoform was classified according to its splice pattern and supporting evidence. We found that one-third of the alternative transcripts examined contain premature termination codons, and most persist even after rigorous filtering by multiple methods. These transcripts are apparent targets of nonsense-mediated mRNA decay (NMD), a surveillance mechanism that selectively degrades nonsense mRNAs. Several of these transcripts are from genes for which alternative splicing is known to regulate protein expression by generating alternate isoforms that are differentially subjected to NMD. We propose that regulated unproductive splicing and translation (RUST), through the coupling of alternative splicing and NMD, may be a pervasive, underappreciated means of regulating protein expression.

Alternative Splicing↗

ASTRAL compendium enhancements.

The ASTRAL compendium provides several databases and tools to aid in the analysis of protein structures, particularly through the use of their sequences. It is partially derived from the SCOP database of protein domains, and it includes sequences for each domain as well as other resources useful for studying these sequences and domain structures. Several major improvements have been made to the ASTRAL compendium since its initial release 2 years ago. The number of protein domain sequences included has doubled from 15 190 to 30 867, and additional databases have been added. The Rapid Access Format (RAF) database contains manually curated mappings linking the biological amino acid sequences described in the SEQRES records of PDB entries to the amino acid sequences structurally observed (provided in the ATOM records) in a format designed for rapid access by automated tools. This information is used to derive sequences for protein domains in the SCOP database. In cases where a SCOP domain spans several protein chains, all of which can be traced back to a single genetic source, a 'genetic domain' sequence is created by concatenating the sequences of each chain in the order found in the original gene sequence. Both the original-style library of SCOP sequences and a new library including genetic domain sequences are available. Selected representative subsets of each of these libraries, based on multiple criteria and degrees of similarity, are also included. ASTRAL may be accessed at http://astral.stanford.edu/.

Amino Acid Sequence↗

SCOP database in 2002: refinements accommodate structural genomics.

The SCOP (Structural Classification of Proteins) database is a comprehensive ordering of all proteins of known structure, according to their evolutionary and structural relationships. Protein domains in SCOP are grouped into species and hierarchically classified into families, superfamilies, folds and classes. Recently, we introduced a new set of features with the aim of standardizing access to the database, and providing a solid basis to manage the increasing number of experimental structures expected from structural genomics projects. These features include: a new set of identifiers, which uniquely identify each entry in the hierarchy; a compact representation of protein domain classification; a new set of parseable files, which fully describe all domains in SCOP and the hierarchy itself. These new features are reflected in the ASTRAL compendium. The SCOP search engine has also been updated, and a set of links to external resources added at the level of domain entries. SCOP can be accessed at http://scop.mrc-lmb.cam.ac.uk/scop.

Animals↗

SCOR: a Structural Classification of RNA database.

The Structural Classification of RNA (SCOR) database provides a survey of the three-dimensional motifs contained in 259 NMR and X-ray RNA structures. In one classification, the structures are grouped according to function. The RNA motifs, including internal and external loops, are also organized in a hierarchical classification. The 259 database entries contain 223 internal and 203 external loops; 52 entries consist of fully complementary duplexes. A classification of the well-characterized tertiary interactions found in the larger RNA structures is also included along with examples. The SCOR database is accessible at http://scor.lbl.gov.

Animals↗

Sulfotransferases and sulfatases in mycobacteria.

Analysis of the genomes of M. tuberculosis, M. leprae, M. smegmatis, and M. avium has revealed a large family of genes homologous to known sulfotransferases. Despite reports detailing a suite of sulfated glycolipids in many mycobacteria, a corresponding family of sulfotransferase genes remains uncharacterized. Here, a sequence-based analysis of newly discovered mycobacterial sulfotransferase genes, named stf1-stf10, is presented. Interestingly, two sulfotransferase genes are highly similar to mammalian sulfotransferases, increasing the list of mycobacterial eukaryotic-like protein families. The sulfotransferases join an equally complex family of mycobacterial sulfatases: a large family of sulfatase genes has been found in all of the mycobacterial genomes examined. As sulfated molecules are common mediators of cell-cell interactions, the sulfotransferases and sulfatases may be involved in regulating host-pathogen interactions.

Amino Acid Sequence↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗