Search PubMedSearch

Biomedical subjects

P Bucher

Publications and source records attributed to P Bucher.

At least 19 recordsLinked to original sources

Mouse interleukin-2 receptor alpha gene expression. Interleukin-1 and interleukin-2 control transcription via distinct cis-acting elements.

We have shown that interleukin-1 (IL-1) and IL-2 control IL-2 receptor alpha (IL-2R alpha) gene transcription in CD4-CD8- murine T lymphocyte precursors. Here we map the cis-acting elements that mediate interleukin responsiveness of the mouse IL-2R alpha gene using a thymic lymphoma-derived hybridoma (PC60). The transcriptional response of the IL-2R alpha gene to stimulation by IL-1 + IL-2 is biphasic. IL-1 induces a rapid, protein synthesis-independent appearance of IL-2R alpha mRNA that is blocked by inhibitors of NF-kappa B activation. It also primes cells to become IL-2 responsive and thereby prepares the second phase, in which IL-2 induces a 100-fold further increase in IL-2R alpha transcripts. Transient transfection experiments show that several elements in the promoter-proximal region of the IL-2R alpha gene contribute to IL-1 responsiveness, most importantly an NF-kappa B site conserved in the human and mouse gene. IL-2 responsiveness, on the other hand, depends on a 78-nucleotide segment 1.3 kilobases upstream of the major transcription start site. This segment functions as an IL-2-inducible enhancer and lies within a region that becomes DNase I hypersensitive in normal T cells in which IL-2R alpha expression has been induced. IL-2 responsiveness requires three distinct elements within the enhancer. Two of these are potential binding sites for STAT proteins.

Animals

The rsp5-domain is shared by proteins of diverse functions.

A novel, unusually small, and highly conserved domain of modular intracellular proteins is described. The domain was first recognized as three repeats in the yeast rsp5 gene product and named thereafter. The rsp5 protein is thought to interact with nuclear proteins but also contains a C2 domain typical for cytoplasmic proteins. Further analyses revealed several additional occurrences of this domain in diverse protein classes, including cytoplasmic signal transduction proteins, gene products interacting with the transcription machinery, structural proteins like dystrophin, and a putative RNA helicase.

Amino Acid Sequence

Improving the sensitivity of the sequence profile method.

The sequence profile method (Gribskov M, McLachlan AD, Eisenberg D, 1987, Proc Natl Acad Sci USA 84:4355-4358) is a powerful tool to detect distant relationships between amino acid sequences. A profile is a table of position-specific scores and gap penalties, providing a generalized description of a protein motif, which can be used for sequence alignments and database searches instead of an individual sequence. A sequence profile is derived from a multiple sequence alignment. We have found 2 ways to improve the sensitivity of sequence profiles: (1) Sequence weights: Usage of individual weights for each sequence avoids bias toward closely related sequences. These weights are automatically assigned based on the distance of the sequences using a published procedure (Sibbald PR, Argos P, 1990, J Mol Biol 216:813-818). (2) Amino acid substitution table: In addition to the alignment, the construction of a profile also needs an amino acid substitution table. We have found that in some cases a new table, the BLOSUM45 table (Henikoff S, Henikoff JG, 1992, Proc Natl Acad Sci USA 89:10915-10919), is more sensitive than the original Dayhoff table or the modified Dayhoff table used in the current implementation. Profiles derived by the improved method are more sensitive and selective in a number of cases where previous methods have failed to completely separate true members from false positives.

Amino Acid Sequence

A generalized profile syntax for biomolecular sequence motifs and its function in automatic sequence interpretation.

A general syntax for expressing biomolecular sequence motifs is described, which will be used in future releases of the PROSITE data bank and in a similar collection of nucleic acid sequence motifs currently under development. The central part of the syntax is a regular structure which can be viewed as a generalization of the profiles introduced by Gribskov and coworkers. Accessory features implement specific motif search strategies and provide information helpful for the interpretation of predicted matches. Two contrasting examples, representing E. coli promoters and SH3 domains respectively, are shown to demonstrate the versatility of the syntax, and its compatibility with diverse motif search methods. It is argued, that a comprehensive machine-readable motif collection based on the new syntax, in conjunction with a standard search program, can serve as a general-purpose sequence interpretation and function prediction tool.

Animals

PROSITE: recent developments.

PROSITE is a compilation of sites and patterns found in protein sequences; it can be used as a method of determining the function of uncharacterized proteins translated from genomic or cDNA sequences.

Amino Acid Sequence

Correlation analysis of amino acid usage in protein classes.

We present a comparative study of residue usage correlations of various organism protein sets of diverse phylogenetic species and of open reading frames of several large human viral genomes. Our correlation analysis reveals three major tendencies: (i) charge compensation reflected by the high correlation of basic with acidic residues; (ii) the positive correlations of functionally and structurally similar amino acids including many pairs of hydrophobic amino acids, all pairs of aromatic amino acids, the anionic pair (glutamate and aspartate), but not the cationic pair (lysine and arginine), moderately the hydroxyl pair (serine and threonine), the small amino acids (glycine and alanine), and many (but not all) of those having high values in the Dayhoff substitutability matrix (characteristics such as amino acid polarity or codon usage agreement, except for the wobble position, do not necessarily imply significant positive correlations); (iii) a widespread negative correlation of the aggregate strong codon group amino acids (Ala, Gly, Pro) versus the weak codon group amino acids (Lys, Ile, Tyr, Asn, Phe). Discussion and speculations relate amino acid usage correlations to protein function/structure, cellular localization, proximity in amino acid biosynthetic pathways, amino acid relative abundances, tRNA and aminoacyl synthetase availabilities, and evolutionary processes.

Amino Acids

Methods and algorithms for statistical analysis of protein sequences.

We describe several protein sequence statistics designed to evaluate distinctive attributes of residue content and arrangement in primary structure. Considered are global compositional biases, local clustering of different residue types (e.g., charged residues, hydrophobic residues, Ser/Thr), long runs of charged or uncharged residues, periodic patterns, counts and distribution of homooligopeptides, and unusual spacings between particular residue types. The computer program SAPS (statistical analysis of protein sequences) calculates all the statistics for any individual protein sequence input and is available for the UNIX environment through electronic mail on request to V.B. (volker/genomic@stanford.edu).

Algorithms

Significant similarity and dissimilarity in homologous proteins.

Common practice emphasizes significant sequence similarities between different members of protein families. These similarities presumably reflect on evolutionary conservation of structurally and functionally essential residues. The nonconserved regions, on the other hand, may be either selectively neutral or differentiated. We propose several distributional sequence statistics (e.g., clustering of charged residues, compositional biases, and repetitive patterns) as indicators of differentiation events. These ideas are illustrated with various examples, including comparisons among G protein-coupled receptors, herpesvirus proteins, and GTPase-activating proteins.

GTP-Binding Proteins

Quantile distributions of amino acid usage in protein classes.

A comparative study of the compositional properties of various protein sets from both cellular and viral organisms is presented. Invariants and contrasts of amino acid usages have been discerned for different protein function classes and for different species using robust statistical methods based on quantile distributions and stochastic ordering relationships. In addition, a quantitative criterion to assess amino acid compositional extremes relative to a reference protein set is proposed and applied. Invariants of amino acid usage relate mainly to the central range of quantile distributions, whereas contrasts occur mainly in the tails of the distributions, especially contrasts between eukaryote and prokaryote species. Influences from genomic constraint are evident, for example, in the arginine:lysine ratios and the usage frequencies of residues encoded by G + C-rich versus A + T-rich codon types. The structurally similar amino acids, glutamate versus aspartate and phenylalanine versus tyrosine, show stochastic dominance relationships for most species protein sets favoring glutamate and phenylalanine respectively. The quantile distribution of hydrophobic amino acid usages in prokaryote data dominates the corresponding quantile distribution in human data. In contrast, glutamate, cysteine, proline and serine usages in human proteins dominate the corresponding quantile distributions in Escherichia coli. E. coli dominates human in the use of basic residues, but no dominance ordering applies to acidic residues. The discussion centers on commonalities and anomalies of the amino acid compositional spectrum in relation to species, function, cellular localization, biochemical and steric attributes, complexity of the amino acid biosynthetic pathway, amino acid relative abundances and founder effects.

Amino Acids

Characterization of a new tissue-specific transcription factor binding to the simian virus 40 enhancer TC-II (NF-kappa B) element.

We have biochemically and functionally characterized a new transcription factor, NP-TCII, which is present in nuclei from unstimulated T and B lymphocytes but is not found in nonhematopoietic cells. This factor has a DNA-binding specificity similar to that of NF-kappa B but is unrelated to this or other Rel proteins by functional and biochemical criteria. It can also be distinguished from other previously described lymphocyte-specific DNA-binding proteins.

Animals

Evidence for selective evolution in codon usage in conserved amino acid segments of human alphaherpesvirus proteins.

The genomes of human viruses herpes simplex 1 (HSV1) and varicella zoster (VZV), although similar in biology, largely concordant in gene order, and identical in many amino acid segments, differ widely in their genomic G + C (abbreviated S) content, which is high in HSV1 (68%) and low in VZV (46%). This paper analyzes several striking codon usage contrasts. The S difference in coding regions is dramatically large in codon site 3, S3, about 42%. The large difference in S3 is maintained at the same level in a subset of closely similar genes and even in corresponding identical amino acid blocks. A similar difference in S levels in silent site 1 (S1) is found in leucine and arginine. The difference in S3 levels occurs in every gene and in every multicodon amino acid form. The S difference also exists in amino acid usage, with HSV1 using significantly more codon types SSN, while VZV uses more codon types WWN (where W stands for A or T). The nonoverlapping and narrow histograms of S3 gene frequencies in both viruses suggest that the difference has arisen and been maintained by a process of selective rather than nonselective effects. This is in sharp contrast to the relatively large variance seen for highly similar genes in the human versus yeast analysis. Interpretations and hypotheses to explain the HSV1 vs VZV codon usage disparity relate to virus-host interactions, to the role of viral genes in DNA metabolism, to availability of molecular resources (molecular Gause exclusion principle), and to differences in genomic structure.

Amino Acids

Occurrence of oligopurine.oligopyrimidine tracts in eukaryotic and prokaryotic genes.

A program to analyse the length and frequency distribution of specific base tracts in genomic sequences is described. The frequency of oligopurine.oligopyrimidine tracts (R.Y. tracts) in a data base of 163 transcribed genes is analysed and compared. The complete genomes of SV40 virus, N. tobacum chloroplast, yeast 2 micron plasmid, bacteriophage lambda, plasmid pBR322 and the E. coli lac operon are also analyzed. A highly significant overrepresentation of oligopurine and oligopyrimidine tracts is observed in all eukaryotic genes examined, as well as in the chloroplast genome. The overrepresentation is evident in all gene subregions of the chloroplast, in the following order: intergenic regions, 3' downstream and 5' upstream (promoter), 5' and 3' untranslated, introns and coding regions. In genes coding for basic proteins, oligopurine rather than oligopyrimidine tracts are found on the coding stand. In prokaryotic genes only the longest R.Y. tracts (greater than or equal to 12) are found in excess, and are concentrated near regulatory regions. While a structural role for R.Y. tracts is most likely in intergenic regions, a functional role, as initiation sites for strand separation, is proposed for regulatory gene regions.

Base Composition

Weight matrix descriptions of four eukaryotic RNA polymerase II promoter elements derived from 502 unrelated promoter sequences.

Optimized weight matrices defining four major eukaryotic promoter elements, the TATA-box, cap signal, CCAAT-, and GC-box, are presented; they were derived by comparative sequence analysis of 502 unrelated RNA polymerase II promoter regions. The new TATA-box and cap signal descriptions differ in several respects from the only hitherto available base frequency Tables. The CCAAT-box matrix, obtained with no prior assumption but CCAAT being the core of the motif, reflects precisely the sequence specificity of the recently discovered nuclear factor NY-I/CP1 but does not include typical recognition sequences of two other purported CCAAT-binding proteins, CTF and CBP. The GC-box description is longer than the previously proposed consensus sequences but is consistent with Sp1 protein-DNA binding data. The notion of a CACCC element distinct from the GC-box seems not to be justified any longer in view of the new weight matrix. Unlike the two fixed-distance elements, neither the CCAAT- nor the GC-box occurs at significantly high frequency in the upstream regions of non-vertebrate genes. Preliminary attempts to predict promoters with the aid of the new signal descriptions were unexpectedly successful. The new TATA-box matrix locates eukaryotic transcription initiation sites as reliably as do the best currently available methods to map Escherichia coli promoters. This analysis was made possible by the recently established Eukaryotic Promoter Database (EPD) of the EMBL Nucleotide Sequence Data Library. In order to derive the weight matrices, a novel algorithm has been devised that is generally applicable to sequence motifs positionally correlated with a biologically defined position in the sequences. The signal must be sufficiently over-represented in a particular region relative to the given site, but need not be present in all members of the input sequence collection. The algorithm iteratively redefines the set of putative motif representatives from which a weight matrix is derived, so as to maximize a quantitative measure of local over-representation, an optimization criterion that naturally combines structural and positional constancy. A comprehensive description of the technique is presented in Methods and Data.

Algorithms

CCAAT box revisited: bidirectionality, location and context.

The so-called CCAAT box is believed to be a major promoter element of higher eukaryotes though it is ill-defined being deduced from very limited sequence data. The comprehensive computer analysis of an unbiased set of 168 promoters presented here removes several of the persisting uncertainties. In particular, it delineates the region of preferential occurrence of the CCAAT element to -110 to -50 relative to the initiation site, suggests that integrity of this pentamer is essential, and confirms bidirectionality as a general property of this element. Within the above region the signal is found to occur in a specific sequence context which is an important supplement to its description.

Animal Population Groups