Search PubMed⌕ Search

Biomedical subjects

Michael Y Galperin

Publications and source records attributed to Michael Y Galperin.

At least 19 recordsLinked to original sources

ComFB, a widespread family of c-di-NMP receptor proteins.

Cyclic dimeric-GMP (c-di-GMP) is a ubiquitous bacterial second messenger that regulates a variety of cellular processes, including motility, biofilm formation, secretion, cell cycle progression, and development, and also contributes to the virulence of many bacterial pathogens. While the genes encoding c-di-GMP cyclases and hydrolases are readily identifiable in microbial genomes, known c-di-GMP receptor domains are quite few, with only PilZ and MshEN broadly distributed across bacterial phyla. Recently, a new c-di-GMP receptor, named CdgR or ComFB, has been identified in cyanobacteria and shown to regulate cell size and natural competence. We demonstrated that CdgR proteins exhibit sequence and structural similarity to the Bacillus subtilis late competence development protein ComFB, a conserved protein of unknown function associated with bacterial competence. This prompted us to hypothesize that ComFB and ComFB-like proteins could also serve as c-di-GMP receptors. Here, we comprehensively investigated the ComFB protein family and demonstrated that ComFB proteins are evolutionarily widespread among bacteria and function as a novel family of c-di-GMP receptors. We showed that ComFB proteins from Gram-positive bacteria (B. subtilis, Thermoanaerobacter brockii) and Gram-negative pathogens (Vibrio cholerae, Treponema denticola) bind c-di-GMP with high affinity. Several ComFB proteins also bind cyclic di-adenosine monophosphate (c-di-AMP), suggesting that ComFB represents a widely distributed bacterial protein family with dual specificity for c-di-GMP and c-di-AMP. Our physiological studies further showed that ComFB plays vital roles in controlling motility in a c-di-GMP-dependent manner in two phylogenetically distant bacteria, B. subtilis and the gram-negative Shewanella oneidensis, attesting to the biological relevance of ComFB as a c-di-GMP binding protein.

Bacterial Proteins↗

Genome-based identification and characterization of a putative mucin-binding protein from the surface of Streptococcus pneumoniae.

Streptococcus pneumoniae open reading frame SP1492 encodes a surface protein that contains a novel conserved domain similar to the repeated fragments of mucin-binding proteins from lactobacilli and lactococci. To investigate the functional role(s) of this protein and its potential adhesive properties, the surface-exposed region of SP1492 was expressed in Escherichia coli, purified to homogeneity, and partially characterized by biophysical and immunological methods. Circular dichroism and sedimentation measurements confirmed that SP1492 is an all-beta protein that exists in solution as a monomer. The SP1492 protein has been shown to be expressed by S. pneumoniae and was experimentally localized to its surface. The protein functional domain binds to mucins II and III from porcine stomach and to purified submaxillary bovine gland mucin. It appears to be one of the very few unambiguous pneumococcal adhesin molecules known to date. A hypothetical model constructed by ab initio techniques predicts a novel beta-sandwich protein structure.

Amino Acid Sequence↗

The Molecular Biology Database Collection: 2007 update.

The NAR online Molecular Biology Database Collection is a public resource that contains links to the databases described in this issue of Nucleic Acids Research, previous NAR database issues, as well as a selection of other molecular biology databases that are freely available on the web and might be useful to the molecular biologist. The 2007 update includes 968 databases, 110 more than the previous one. Many databases that have been described in earlier issues of NAR come with updated summaries, which reflect recent progress and, in some instances, an expanded scope of these databases. The complete database list and summaries are available online on the Nucleic Acids Research web site http://nar.oxfordjournals.org/.

Databases, Genetic↗

Sentra: a database of signal transduction proteins for comparative genome analysis.

Sentra (http://compbio.mcs.anl.gov/sentra), a database of signal transduction proteins encoded in completely sequenced prokaryotic genomes, has been updated to reflect recent advances in understanding signal transduction events on a whole-genome scale. Sentra consists of two principal components, a manually curated list of signal transduction proteins in 202 completely sequenced prokaryotic genomes and an automatically generated listing of predicted signaling proteins in 235 sequenced genomes that are awaiting manual curation. In addition to two-component histidine kinases and response regulators, the database now lists manually curated Ser/Thr/Tyr protein kinases and protein phosphatases, as well as adenylate and diguanylate cyclases and c-di-GMP phosphodiesterases, as defined in several recent reviews. All entries in Sentra are extensively annotated with relevant information from public databases (e.g. UniProt, KEGG, PDB and NCBI). Sentra's infrastructure was redesigned to support interactive cross-genome comparisons of signal transduction capabilities of prokaryotic organisms from a taxonomic and phenotypic perspective and in the framework of signal transduction pathways from KEGG. Sentra leverages the PUMA2 system to support interactive analysis and annotation of signal transduction proteins by the users.

Archaeal Proteins↗

New metrics for comparative genomics.

The availability of genome sequences from a variety of organisms presents an opportunity to apply this sequence information to solving the key problems of molecular biology. One of the principal roadblocks on this path is the lack of appropriate descriptors and metrics that could succinctly represent the new knowledge stemming from the genomic data. Several new metrics have recently been used in comparative genome analysis, yet challenges remain in finding an appropriate language for the emerging discipline of systems biology.

Animals↗

The cyanobacterial genome core and the origin of photosynthesis.

Comparative analysis of 15 complete cyanobacterial genome sequences, including "near minimal" genomes of five strains of Prochlorococcus spp., revealed 1,054 protein families [core cyanobacterial clusters of orthologous groups of proteins (core CyOGs)] encoded in at least 14 of them. The majority of the core CyOGs are involved in central cellular functions that are shared with other bacteria; 50 core CyOGs are specific for cyanobacteria, whereas 84 are exclusively shared by cyanobacteria and plants and/or other plastid-carrying eukaryotes, such as diatoms or apicomplexans. The latter group includes 35 families of uncharacterized proteins, which could also be involved in photosynthesis. Only a few components of cyanobacterial photosynthetic machinery are represented in the genomes of the anoxygenic phototrophic bacteria Chlorobium tepidum, Rhodopseudomonas palustris, Chloroflexus aurantiacus, or Heliobacillus mobilis. These observations, coupled with recent geological data on the properties of the ancient phototrophs, suggest that photosynthesis originated in the cyanobacterial lineage under the selective pressures of UV light and depletion of electron donors. We propose that the first phototrophs were anaerobic ancestors of cyanobacteria ("procyanobacteria") that conducted anoxygenic photosynthesis using a photosystem I-like reaction center, somewhat similar to the heterocysts of modern filamentous cyanobacteria. From procyanobacteria, photosynthesis spread to other phyla by way of lateral gene transfer.

Bacterial Proteins↗

Cyanobacterial response regulator PatA contains a conserved N-terminal domain (PATAN) with an alpha-helical insertion.

The cyanobacterium Anabaena (Nostoc) PCC 7120 responds to starvation for nitrogen compounds by differentiating approximately every 10th cell in the filament into nitrogen-fixing cells called heterocysts. Heterocyst formation is subject to complex regulation, which involves an unusual response regulator PatA that contains a CheY-like phosphoacceptor (receiver, REC) domain at its C-terminus. PatA-like response regulators are widespread in cyanobacteria; one of them regulates phototaxis in Synechocystis PCC 6803. Sequence analysis of PatA revealed, in addition to the REC domain, a previously undetected, conserved domain, which we named PATAN (after PatA N-terminus), and a potential helix-turn-helix (HTH) domain. PATAN domains are encoded in a variety of environmental bacteria and archaea, often in several copies per genome, and are typically associated with REC, Roadblock and other signal transduction domains, or with DNA-binding HTH domains. Many PATAN domains contain insertions of a small additional domain, termed alpha-clip, which is predicted to form a four-helix bundle. PATAN domains appear to participate in protein-protein interactions that regulate gliding motility and processes of cell development and differentiation in cyanobacteria and some proteobacteria, such as Myxococcus xanthus and Geobacter sulfurreducens.

Amino Acid Sequence↗

The Molecular Biology Database Collection: 2006 update.

The NAR Molecular Biology Database Collection is a public online resource that contains links to all databases described in this issue of Nucleic Acids Research. In addition, this collection lists databases that have been featured in previous issues of NAR, as well as selected other databases that are freely available to the public and may be useful to the molecular biologist. The 2006 update includes 858 databases, 139 more than the previous one. The databases come with brief summaries, many of which have been updated recently. Each database is assigned a stable accession number that does not change if the database moves to a new location and its URL, authors' names or the contact person address are updated. The complete database list and summaries are available online at the Nucleic Acids Research website http://nar.oxfordjournals.org/.

Databases, Genetic↗

House cleaning, a part of good housekeeping.

Cellular metabolism constantly generates by-products that are wasteful or even harmful. Such compounds are excreted from the cell or are removed through hydrolysis to normal cellular metabolites by various 'house-cleaning' enzymes. Some of the most important contaminants are non-canonical nucleoside triphosphates (NTPs) whose incorporation into the nascent DNA leads to increased mutagenesis and DNA damage. Enzymes intercepting abnormal NTPs from incorporation by DNA polymerases work in parallel with DNA repair enzymes that remove lesions produced by modified nucleotides. House-cleaning NTP pyrophosphatases targeting non-canonical NTPs belong to at least four structural superfamilies: MutT-related (Nudix) hydrolases, dUTPase, ITPase (Maf/HAM1) and all-alpha NTP pyrophosphatases (MazG). These enzymes have high affinity (Km's in the micromolar range) for their natural substrates (8-oxo-dGTP, dUTP, dITP, 2-oxo-dATP), which allows them to select these substrates from a mixture containing a approximately 1000-fold excess of canonical NTPs. To date, many house-cleaning NTPases have been identified only on the basis of their side activity towards canonical NTPs and NDP derivatives. Integration of growing structural and biochemical data on these superfamilies suggests that their new family members cleanse the nucleotide pool of the products of oxidative damage and inappropriate methylation. House-cleaning enzymes, such as 6-phosphogluconolactonase, are also part of normal intermediary metabolism. Genomic data suggest that house-cleaning systems are more abundant than previously thought and include numerous analogous enzymes with overlapping functions. We discuss the structural diversity of these enzymes, their phylogenetic distribution, substrate specificity and the problem of identifying their true substrates.

Binding Sites↗

Structural classification of bacterial response regulators: diversity of output domains and domain combinations.

CheY-like phosphoacceptor (or receiver [REC]) domain is a common module in a variety of response regulators of the bacterial signal transduction systems. In this work, 4,610 response regulators, encoded in complete genomes of 200 bacterial and archaeal species, were identified and classified by their domain architectures. Previously uncharacterized output domains were analyzed and, in some cases, assigned to known domain families. Transcriptional regulators of the OmpR, NarL, and NtrC families were found to comprise almost 60% of all response regulators; transcriptional regulators with other DNA-binding domains (LytTR, AraC, Spo0A, Fis, YcbB, RpoE, and MerR) account for an additional 6%. The remaining one-third is represented by the stand-alone REC domain (approximately 14%) and its combinations with a variety of enzymatic (GGDEF, EAL, HD-GYP, CheB, CheC, PP2C, and HisK), RNA-binding (ANTAR and CsrA), protein- or ligand-binding (PAS, GAF, TPR, CAP_ED, and HPt) domains, or newly described domains of unknown function. The diversity of domain architectures and the abundance of alternative domain combinations suggest that fusions between the REC domain and various output domains is a widespread evolutionary mechanism that allows bacterial cells to regulate transcription, enzyme activity, and/or protein-protein interactions in response to environmental challenges. The complete list of response regulators encoded in each of the 200 analyzed genomes is available online at http://www.ncbi.nlm.nih.gov/Complete_Genomes/RRcensus.html.

Amino Acid Sequence↗

PilZ domain is part of the bacterial c-di-GMP binding protein.

Recent studies identified c-di-GMP as a universal bacterial secondary messenger regulating biofilm formation, motility, production of extracellular polysaccharide and multicellular behavior in diverse bacteria. However, except for cellulose synthase, no protein has been shown to bind c-di-GMP and the targets for c-di-GMP action remain unknown. Here we report identification of the PilZ ("pills") domain (Pfam domain PF07238) in the sequences of bacterial cellulose synthases, alginate biosynthesis protein Alg44, proteins of enterobacterial YcgR and firmicute YpfA families, and other proteins encoded in bacterial genomes and present evidence indicating that this domain is (part of) the long-sought c-di-GMP-binding protein. Association of the PilZ domain with a variety of other domains, including likely components of bacterial multidrug secretion system, could provide clues to multiple functions of the c-di-GMP in bacterial pathogenesis and cell development.

Alginates↗