Search PubMedSearch

Biomedical subjects

M S Boguski

Publications and source records attributed to M S Boguski.

At least 19 recordsLinked to original sources

Threading analysis suggests that the obese gene product may be a helical cytokine.

The ob gene encodes a protein that, in mutant form, is associated with obesity and type II diabetes in mice. Sequence analysis has revealed no similarities to other proteins, however, and no clues as to possible functions. The possibility nonetheless remains that ob is functionally or ancestrally related to other proteins, whose sequences are divergent to the point that only a comparison of three-dimensional structures might detect relationship. To explore this possibility, we conduct a 'threading' search of a 3-dimensional structure database, to determine whether the ob protein might adopt a fold similar to any known structure. This search reveals that the ob sequence is compatible, at a significance level of P < 0.05, with structures from the family of helical cytokines that includes interleukin-2 and growth hormone. A structural model of ob based upon these results is physically and biologically plausible and leads to testable predictions, including the prediction that ob may activate the JAK-STAT pathway, via binding to a receptor resembling those of the cytokine family.

Amino Acid Sequence

A novel RING finger protein interacts with the cytoplasmic domain of CD40.

CD40 is a member of the tumor necrosis factor receptor family and, like other members, it appears to possess no intrinsic signaling capacity (e.g. kinase activity), suggesting that signal transduction is likely mediated by associating molecules. To identify such molecules, we have utilized the yeast two-hybrid system to clone cDNAs encoding proteins that bind the CD40 cytoplasmic domain. One such interacting protein, designated CD40-binding protein, has a N-terminal RING finger motif that is found in a number of DNA-binding proteins, including the V(D)J recombination activating gene RAG1. In addition, it contains a prominent central coiled-coil segment that may allow homo- or hetero-oligomerization. The C terminus possesses substantial homology to the tumor necrosis factor receptor-associated factor (TRAF) domain that is found in two proteins (TRAF1 and TRAF2) that associate with the cytoplasmic domain of the related 75-kDa tumor necrosis factor receptor. This is the first identification of a molecule that interacts with CD40 and whose sequence suggests a potential role in signaling.

Amino Acid Sequence

Bioinformatics.

Computer databases, networks and software tools are essential materials and methods for biomedical research and are involved in almost every aspect of disease gene mapping and positional cloning. Public databases of DNA and protein sequences and genetic and physical map information are increasing rapidly in size and complexity and are also improving in quality, comprehensiveness, interoperability and access. A new generation of software tools for navigating through the biomedical literature has become available. Programs for sequence homology searching and genetic map construction have become more sophisticated, yet easier to use. Global computer networks are bringing state-of-the-art capabilities to all.

Chromosome Mapping

Issues in searching molecular sequence databases.

Sequence similarity search programs are versatile tools for the molecular biologist, frequently able to identify possible DNA coding regions and to provide clues to gene and protein structure and function. While much attention had been paid to the precise algorithms these programs employ and to their relative speeds, there is a constellation of associated issues that are equally important to realize the full potential of these methods. Here, we consider a number of these issues, including the choice of scoring systems, the statistical significance of alignments, the masking of uninformative or potentially confounding sequence regions, the nature and extent of sequence redundancy in the databases and network access to similarity search services.

Algorithms

Genes conserved in yeast and humans.

Evolutionary conservation of homologous gene products from distantly related organisms provides an information resource of great value for elucidating protein structure and function. Sequence similarities also serve as molecular cross-references between diverse organisms that offer different, or complementary, experimental approaches for analyzing gene expression and biochemistry in normal and abnormal states. There are now countless examples of information about a protein from one species contributing to the understanding of biological phenomena or disease in another species. Such connections are often unanticipated and surprising, but there is an opportunity to make them more systematically as concerted genome sequencing projects progress. In the present review we focus on connections between yeast and human proteins and their functional implications. We present several 'case studies' as well as survey results derived from comprehensive sequence comparisons among all yeast and human proteins currently present in the public databases.

Conserved Sequence

Proteins regulating Ras and its relatives.

GTPases of the Ras superfamily regulate many aspects of cell growth, differentiation and action. Their functions depend on their ability to alternate between inactive and active forms, and on their cellular localization. Numerous proteins affecting the GTPase activity, nucleotide exchange rates and membrane localization of Ras superfamily members have now been identified. Many of these proteins are much larger and more complex than their targets, containing multiple domains capable of interacting with an intricate network of cellular enzymes and structures.

Amino Acid Sequence

Linking yeast genetics to mammalian genomes: identification and mapping of the human homolog of CDC27 via the expressed sequence tag (EST) data base.

We describe a strategy for quickly identifying and positionally mapping human homologs of yeast genes to cross-reference the biological and genetic information known about yeast genes to mammalian chromosomal maps. Optimized computer search methods have been developed to scan the rapidly expanding expressed sequence tag (EST) data base to find human open reading frames related to yeast protein sequence queries. These methods take advantage of the newly developed BLOSUM scoring matrices and the query masking function SEG. The corresponding human cDNA is then used to obtain a high-resolution map position on human and mouse chromosomes, providing the links between yeast genetic analysis and mapped mammalian loci. By using these methods, a human homolog of Saccharomyces cerevisiae CDC27 has been identified and mapped to human chromosome 17 and mouse chromosome 11 between the Pkca and Erbb-2 genes. Human CDC27 encodes an 823-aa protein with global similarity to its fungal homologs CDC27, nuc2+, and BimA. Comprehensive cross-referencing of genes and mutant phenotypes described in humans, mice, and yeast should accelerate the study of normal eukaryotic biology and human disease states.

Amino Acid Sequence

Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.

A wealth of protein and DNA sequence data is being generated by genome projects and other sequencing efforts. A crucial barrier to deciphering these sequences and understanding the relations among them is the difficulty of detecting subtle local residue patterns common to multiple sequences. Such patterns frequently reflect similar molecular structures and biological properties. A mathematical definition of this "local multiple alignment" problem suitable for full computer automation has been used to develop a new and sensitive algorithm, based on the statistical method of iterative sampling. This algorithm finds an optimized local alignment model for N sequences in N-linear time, requiring only seconds on current workstations, and allows the simultaneous detection and optimization of multiple patterns and pattern repeats. The method is illustrated as applied to helix-turn-helix proteins, lipocalins, and prenyltransferases.

Algorithms

Comparative analysis of the beta transducin family with identification of several new members including PWP1, a nonessential gene of Saccharomyces cerevisiae that is divergently transcribed from NMT1.

While investigating the expression of the Saccharomyces cerevisiae myristoyl-CoA:protein N-myristoyltransferase gene (NMT: E.C. 2.3.1.97) by Northern blot analysis, we observed another RNA transcript whose expression resembled that of NMT1 during meiosis and was derived from a gene located less than 1 kb immediately upstream of NMT1. This new gene, designated PWP1 (for periodic tryptophan protein), is divergently transcribed from NMT1 and encodes a 576-residue protein. Null mutants of PWP1 are viable, but their growth is severely retarded and steady-state levels of several cellular proteins (including at least two proteins that label with exogenous [3H]myristic acid) are drastically reduced. New methods for database searching and assessing the statistical significance of sequence similarities identify PWP1 as a member of the beta-transducin protein superfamily. Two other previously unrecognized beta-transducin-like proteins (S. cerevisiae MAK11 and D. discoideum AAC3) were also identified, and an unexpectedly high degree of sequence homology was found between a Chlamydomonas beta-like polypeptide and the C12.3 gene of chickens. A systematic and quantitative comparative analysis resulted in classifying all beta-transducin-like sequences into 11 nonorthologous families. Based on specific sequence attributes, however, not all beta-transducin-like sequences are expected to be functionally similar, and quantitative criteria for inferring functional analogies are discussed. Possible roles of repetitive tryptophan residues in proteins are also considered.

Acyltransferases

Computational sequence analysis revisited: new databases, software tools, and the research opportunities they engender.

The increasing quantity and complexity of sequences and structural data for proteins and nucleic acids create both problems and opportunities for biomedical researchers. Fortunately, a new generation of practical computer tools for data analysis and integrated information retrieval is emerging. Recent developments in fast database searching, multiple sequence alignment, and molecular modeling are discussed and windows-based, mouse-driven software for CD-ROM and network information retrieval are described. Each method is illustrated with a practical example pertinent to lipid research. In particular, the connection among cholesteryl ester transfer protein, bactericidal permeability-increasing protein, and lipopolysaccharide-binding proteins is determined; novel repetitive sequence motifs in mammalian farnesyltransferase subunits and related yeast prenyltransferases are derived; biochemical insights from a three-dimensional model of human apolipoprotein D based on two insect lipocalins are discussed; the relationship between apolipoprotein D and gross cystic disease fluid protein from human breast is reviewed; and prospects for modeling apolipoprotein E-related proteins are described. In addition, information on a number of general and special-purpose sequence, motif, and structural databases is included.

Amino Acid Sequence