Search PubMedSearch

Biomedical subjects

V Brendel

Publications and source records attributed to V Brendel.

At least 19 recordsLinked to original sources

Chance and statistical significance in protein and DNA sequence analysis.

Statistical approaches help in the determination of significant configurations in protein and nucleic acid sequence data. Three recent statistical methods are discussed: (i) score-based sequence analysis that provides a means for characterizing anomalies in local sequence text and for evaluating sequence comparisons; (ii) quantile distributions of amino acid usage that reveal general compositional biases in proteins and evolutionary relations; and (iii) r-scan statistics that can be applied to the analysis of spacings of sequence markers.

Amino Acid Sequence

Methods and algorithms for statistical analysis of protein sequences.

We describe several protein sequence statistics designed to evaluate distinctive attributes of residue content and arrangement in primary structure. Considered are global compositional biases, local clustering of different residue types (e.g., charged residues, hydrophobic residues, Ser/Thr), long runs of charged or uncharged residues, periodic patterns, counts and distribution of homooligopeptides, and unusual spacings between particular residue types. The computer program SAPS (statistical analysis of protein sequences) calculates all the statistics for any individual protein sequence input and is available for the UNIX environment through electronic mail on request to V.B. (volker/genomic@stanford.edu).

Algorithms

Significant similarity and dissimilarity in homologous proteins.

Common practice emphasizes significant sequence similarities between different members of protein families. These similarities presumably reflect on evolutionary conservation of structurally and functionally essential residues. The nonconserved regions, on the other hand, may be either selectively neutral or differentiated. We propose several distributional sequence statistics (e.g., clustering of charged residues, compositional biases, and repetitive patterns) as indicators of differentiation events. These ideas are illustrated with various examples, including comparisons among G protein-coupled receptors, herpesvirus proteins, and GTPase-activating proteins.

GTP-Binding Proteins

Very long charge runs in systemic lupus erythematosus-associated autoantigens.

Systemic lupus erythematosus and other chronic systemic autoimmune diseases are associated with circulating autoantibodies reactive with a limited set of mostly nuclear proteins. Using rigorous statistical methods we have identified segments of highly significant charge concentration in the majority of the characteristic nuclear and cytoplasmic autoantigens. Extremely long runs of charged residues, including some sequences of greater than 20 consecutive charged residues (purely acidic or mixed basic and acidic), occur in about a third of these proteins, whereas equivalent runs are found in less than 3% of other mammalian proteins. The other sequences have less extreme charge clusters, the type and location of which are often conserved between several otherwise nonsimilar antigens. We propose that supercharged surfaces render the targeted host proteins strongly immunogenic and that antinuclear antibody profiles might result from chronic exposure to intracellular contents, possibly in conjunction with crossreactive viral products. The limited number of potential systemic autoantigens may partly be due to the rarity of requisite charge properties.

Amino Acid Sequence

Charge configurations in oncogene products and transforming proteins.

Statistically significant charge clusters are of infrequent occurrence in all kinds of proteins. In the six standard classes of proto-oncogene products, all of the nuclear class contain a significant charge cluster and several, but not all, of the transmembrane class do, whereas significant charge clusters or patterns are not found in protooncogenes of primarily cytoplasmic location, nor in membrane-bound (src-like) proto-oncogenes, nor in those of the ras family. Among nuclear oncogene families, such as myc, jun, fos, myb, or ets-related, and among homologous proteins across species, the significant charge clusters are part of the most conserved region. These gene families generally have similar charge distributions embodying a significant charge cluster, not of an invariant sign, preceded by a substantial uncharged stretch of predominantly polar residues. The nuclear transforming proteins p53 and p68 also contain significant charge clusters together with long uncharged segments, suggestive of a modular structure of these proteins. The transmembrane oncogene c-mas contains a mixed charge cluster and c-fms displays an unusual (0, +)7 pattern, in both cases positioned within their intracellular activating domain. Distinctive charge configurations for excreted proto-oncogenes are of a mixed character. Possible functions, mechanisms, and associated experimental procedures for studying proteins with anomalous charge distributions are discussed.

Amino Acid Sequence

Kinetics of complementary RNA-RNA interaction involved in plasmid ColE1 copy number control.

Binding of a small antisense RNA (RNA I) to the primer transcript (RNA II) of plasmid ColE1 inhibits formation of primer for DNA polymerase I-mediated plasmid replication. It is thought that RNA I and RNA II transiently interact via their single-stranded loop regions to form an unstable complex that subsequently converts into a more stable complex by hybridization. Rom (or Rop) protein enhances the inhibitory effect of RNA I on replication by enhancing the binding of the two RNAs. In this paper, we develop a model for the kinetics of the RNA I-RNA II binding reaction, estimate the rate constants, and provide a quantitative description of the effects of Rom protein. We show that the reaction kinetics are consistent with a stepwise binding model in which Rom protein binds to RNA I and RNA II, while the RNAs are held together in a transient complex. Mutations that replace C.G pairs by T.A pairs in the RNA loop regions and thus display weaker hydrogen bonding between the loop regions should be associated with an increased rate of dissociation for the unstable complex. Our model predicts that such destabilization of the loop interactions leads to a greater enhancement in the binding rate by Rom protein. The available data support this prediction.

Bacterial Proteins

A method to identify distinctive charge configurations in protein sequences, with application to human herpesvirus polypeptides.

Charge interactions are of great importance for protein function and structure, and for a variety of cellular and biochemical processes. We present a systematic approach to the detection of distinctive clusters, runs and periodic patterns of charged residues in a protein sequence. Criteria and formulae are set forth to assess statistical significance of these charge configurations. For the 80-odd proteins potentially encoded by the Epstein-Barr virus, only the major nuclear antigens of the latent state and the transactivator of the lytic cycle contain separated charge clusters of opposite sign as well as periodic charge patterns. From our studies of the polypeptides of the human herpesviruses and of a broad collection of human and other viral protein sequences, distinctive charge configurations appear to be associated with viral capsid and core proteins (positive clusters or runs, mostly at the carboxyl terminus), with many viral glycoproteins and membrane-associated proteins (negative charge clusters), and with transactivators and transforming proteins (multiple charge structures). The statistics developed in this paper apply more generally to other than charge properties of a protein and should aid in the evaluation of a large variety of sequence features.

Cytomegalovirus

Association of charge clusters with functional domains of cellular transcription factors.

Using rigorous statistical methods, we have identified and evaluated unusual properties of the distribution of charged residues within the sequences of eukaryotic cellular transcription factors. Virtually all transcription factors, including GAL4, c-Jun, C/EBP, CREB, Oct-1, Oct-2, Sp1, Egr-1, CTF-1, steroid and thyroid hormone receptors, and others, carry one or more highly significant charge clusters. For the most part these clusters (conserved within families of homologous proteins) are of positive net charge but contain also substantial numbers of acidic residues. Predominantly basic charge clusters are often, but not exclusively, associated with DNA-binding domains, and vice versa. Negative charge clusters of note occur only in the yeast protein PHO4 and in the proteins encoded at the Drosophila loci zeste (zeta) and knrl. This dearth of statistically significant negative charge clusters raises questions with respect to the generality of acidic activation domains. A number of sequences (Oct-1, Oct-2, zeste, Dhr23, E75, and knrl) contain multiple charge clusters together with one or more significantly long uncharged regions. The occurrence of multiple charge clusters is a rare phenomenon (found in less than 3% of all proteins, mainly in Drosophila developmental control proteins and in transactivators of eukaryotic DNA viruses). Most of the proteins with zinc-binding "fingers" carry a mixed charge cluster centered at the zinc-finger motif preceded by a long uncharged stretch, suggestive of a modular structure for these proteins.

Animals

Charge configurations in viral proteins.

The spatial distribution of the charged residues of a protein is of interest with respect to potential electrostatic interactions. We have examined the proteins of a large number of representative eukaryotic and prokaryotic viruses for the occurrence of significant clusters, runs, and periodic patterns of charge. Clusters and runs of positive charge are prominent in many capsid and core proteins, whereas surface (glyco)proteins frequently contain a negative charge cluster. Significant charge configurations are abundant in regulatory proteins implicated in transcriptional transactivation and cellular transformation. Proteins with charge structures are much more predominant in animal DNA viruses as compared to animal RNA viruses and prokaryotic viruses. This contrast might reflect the role of protein charge structures in facilitating competitive virus-host interactions involving the cellular transcription, translation, protein sorting, and transport apparatus.

Adenoviridae

Intervening sequences exhibit distinct vocabulary.

Little is known about the origin and function of eukaryotic introns. Application of a novel linguistic approach to the analysis of intervening sequences reveals, however, that they exhibit a specific non-random vocabulary whose major feature is the utilization of mirror-symmetrical words ("mirrorrim"). Introns also manifest a significant tendency to avoid local complementarities. Possible biological implications of the corresponding loop regions in the RNA transcripts are discussed.

Animals

Linguistics of nucleotide sequences: morphology and comparison of vocabularies.

The concept of "words" in continuous languages devoid of blanks is introduced and an operational definition of words given. With this novel concept nucleotide sequences become object for linguistic analysis. The typical word size of the nucleotide language is found to be 3 to 5 (tri- to pentamers). Different genomes have distinct vocabularies. Comparison of these vocabularies can serve as a basis for revealing functional and evolutionary relatedness of sequences.

Bacteriophages

Terminators of transcription with RNA polymerase from Escherichia coli: what they look like and how to find them.

We present here a compilation of prokaryotic transcription terminator sequences (ref. 1-152). The compilation includes 49 independent terminators, 52 speculated independent terminators, 27 sites shown to function in vivo, and some 20 proven or speculated rho-dependent terminators. In addition to the well-known features of independent terminators (dyad symmetry and T-run), two consensus are found: CGGG(C/G) upstream and TCTG downstream of the termination point. A subset of the collection of sequence has been used to construct a computer algorithm to locate independent terminators by sequence analysis.

Algorithms

A computer algorithm for testing potential prokaryotic terminators.

The nucleotide sequences of 30 factor-independent terminators of transcription with RNA polymerase from E. coli have been compiled and analyzed. The standard features - a stretch of thymine residues and a preceding dyad symmetry - are shared by most sequences, but there are striking exceptions which indicate that these features alone are not sufficient to describe these sites. In two thirds of the sequences the 3'-half of the dyad symmetry contains the pentanucleotide CGGG (G/C) or a close derivative; about one third have TCTG or a close derivative just downstream of the termination point. The TCTG -box might be implied in termination of stringently controlled operons of E. coli. An algorithm to locate terminators in templates of known nucleotide sequence has been constructed on the basis of correlation to the distribution of dinucleotides along the aligned signal sequences. The algorithm has been tested on natural sequences of a total length of about 11,500 N. It finds all known independent terminators and only a few other sites, including some of the rho-dependent and putative terminators.

Base Sequence