Search PubMed⌕ Search

Biomedical subjects

Ken Nishikawa

Publications and source records attributed to Ken Nishikawa.

At least 19 recordsLinked to original sources

Gene cluster analysis method identifies horizontally transferred genes with high reliability and indicates that they provide the main mechanism of operon gain in 8 species of gamma-Proteobacteria.

The formation mechanism of operons remains unresolved: operons may form by rearrangements within a genome or by acquisition of genes from other species, that is, horizontal gene transfer (HGT). One hindrance to its elucidation is the unavailability of a method to accurately identify HGT, although it is generally considered to occur. It is critically important first to select horizontally transferred (HT) genes reliably and then to determine the extent to which HGT is involved in operon formation. For this purpose, we considered indels in terms of gene clusters instead of individual genes and chose candidates of HT genes in 8 species of Escherichia, Shigella, and Salmonella based on the minimization of indels. To select a benchmark set of positively HT genes against which we can evaluate the candidate set, we devised another procedure using intergenetic alignments. Comparison with the benchmark set demonstrated the absence of a significant number of false positives in the candidate set, showing the high reliability of the method. Analyses of Escherichia coli K-12 operons revealed that although approximately 20 operons were probably gained from the last common ancestor of the 8 gamma-proteobacteria, deletion of intervening genes accounts for the formation of no operons, whereas horizontal transfer expanded 2 operons and introduced 4 entire operons. Based on these observations and reasoning, we suggest that the main mechanism of operon gain is HGT rather than intragenomic rearrangements. We propose that genes with related essential functions tend to reside in conserved operons, whereas genes in nonconserved operons mostly confer slight advantage to the organisms and frequently undergo horizontal transfer and decay. HT genes constitute at least 5.5% of the genes in the 8 species and approximately 45% of which originate from other gamma-proteobacteria. Genes involved in viral functions and mobile and extrachromosomal element functions are HT more often than expected. This finding indicates frequent mediation of HGT by bacteriophages. On the other hand, not only informational genes (those involved in transcription, translation, and related processes) but also operational genes (those involved in housekeeping) are HT less frequently than expected.

Cluster Analysis↗

CRNPRED: highly accurate prediction of one-dimensional protein structures by large-scale critical random networks.

BACKGROUND: One-dimensional protein structures such as secondary structures or contact numbers are useful for three-dimensional structure prediction and helpful for intuitive understanding of the sequence-structure relationship. Accurate prediction methods will serve as a basis for these and other purposes. RESULTS: We implemented a program CRNPRED which predicts secondary structures, contact numbers and residue-wise contact orders. This program is based on a novel machine learning scheme called critical random networks. Unlike most conventional one-dimensional structure prediction methods which are based on local windows of an amino acid sequence, CRNPRED takes into account the whole sequence. CRNPRED achieves, on average per chain, Q3 = 81% for secondary structure prediction, and correlation coefficients of 0.75 and 0.61 for contact number and residue-wise contact order predictions, respectively. CONCLUSION: CRNPRED will be a useful tool for computational as well as experimental biologists who need accurate one-dimensional protein structure predictions.

Algorithms↗

Human transcription factors contain a high fraction of intrinsically disordered regions essential for transcriptional regulation.

Human transcriptional regulation factors, such as activators, repressors, and enhancer-binding factors are quite different from their prokaryotic counterparts in two respects: the average sequence in human is more than twice as long as that in prokaryotes, while the fraction of sequence aligned to domains of known structure is 31% in human transcription factors (TFs), less than half of that in bacterial TFs (72%). Intrinsically disordered (ID) regions were identified by a disorder-prediction program, and were found to be in good agreement with available experimental data. Analysis of 401 human TFs with experimental evidence from the Swiss-Prot database showed that as high as 49% of the entire sequence of human TFs is occupied by ID regions. More than half of the human TFs consist of a small DNA binding domain (DBD) and long ID regions frequently sandwiching unassigned regions. The remaining TFs have structural domains in addition to DBDs and ID regions. Experimental studies, particularly those with NMR, revealed that the transactivation domains in unbound TFs are usually unstructured, but become structured upon binding to their partners. The sequences of human and mouse TF orthologues are 90.5% identical despite a high incidence of ID regions, probably reflecting important functional roles played by ID regions. In general ID regions occupy a high fraction in TFs of eukaryotes, but not in prokaryotes. Implications of this dichotomy are discussed in connection with their functional roles in transcriptional regulation and evolution.

Animals↗

Stabilization of E. coli Ribonuclease HI by the 'stability profile of mutant protein' (SPMP)-inspired random and non-random mutagenesis.

The change in the structural stability of Escherichia coli ribonuclease HI (RNase HI) due to single amino acid substitutions has been estimated computationally by the stability profile of mutant protein (SPMP) [Ota, M., Kanaya, S. Nishikawa, K., 1995. Desk-top analysis of the structural stability of various point mutations introduced into ribonuclease H. J. Mol. Biol. 248, 733-738]. As well, an effective strategy using random mutagenesis and genetic selection has been developed to obtain E. coli RNase HI mutants with enhanced thermostability [Haruki, M., Noguchi, E., Akasako, A., Oobatake, M., Itaya, M., Kanaya, S., 1994. A novel strategy for stabilization of Escherichia coli ribonuclease HI involving a screen for an intragenic suppressor of carboxyl-terminal deletions. J. Biol. Chem. 269, 26904-26911]. In this study, both methods were combined: random mutations were individually introduced to Lys99-Val101 on the N-terminus of the alpha-helix IV and the preceding beta-turn, where substitutions of other amino acid residues were expected to significantly increase the stability from SPMP, and then followed by genetic selection. Val101 to Ala, Gln, and Arg mutations were selected by genetic selection. The Val101-->Ala mutation increased the thermal stability of E. coli RNase HI by 2.0 degrees C in Tm at pH 5.5, whereas the Val101-->Gln and Val101-->Arg mutations decreased the thermostability. Separately, the Lys99-->Pro and Asn100-->Gly mutations were also introduced directly. The Lys99-->Pro mutation increased the thermostability of E. coli RNase HI by 1.8 degrees C in Tm at pH 5.5, whereas the Asn100-->Gly mutation decreased the thermostability by 17 degrees C. In addition, the Lys99-->Pro mutation altered the dependence of the enzymatic activity on divalent metal ions.

Amino Acid Sequence↗

Genome-wide survey of transcription factors in prokaryotes reveals many bacteria-specific families not found in archaea.

Assignment of all transcription factors (TFs) from genome sequence data is not a straightforward task due to the wide variation in TFs among different species. A DNA binding domain (DBD) and a contiguous non-DBD with a characteristic SCOP or Pfam domain combination are observed in most members of TF families. We found that most of the experimentally verified TFs in prokaryotes are detectable by a combination of SCOP or Pfam domains assigned to DBDs and non-DBDs. Based on this finding, we set up rules to detect TFs and classify them into 52 TF families. Application of the rules to 154 entirely sequenced prokaryotic genomes detected >18,000 TFs classified into families, which have been made publicly available from the 'GTOP_TF' database. Despite the rough proportionality of the number of TFs per genome with genome size, species with reduced genomes, i.e. obligatory parasites and symbionts, have only a few if any TFs, reflecting a nearly complete loss. Also the number of TFs is significantly lower in archaea than in bacteria. In addition, all but 1 of the 19 TF families present in archaea is present in bacteria, whereas 33 TF families are found exclusively in bacteria. This observation indicates that a number of new TF families have evolved in bacteria, making the transcription regulatory system more divergent in bacteria than in archaea.

Algorithms↗

Analysis of amino acid residues involved in catalysis of polyethylene glycol dehydrogenase from Sphingopyxis terrae, using three-dimensional molecular modeling-based kinetic characterization of mutants.

Polyethylene glycol dehydrogenase (PEGDH) from Sphingopyxis terrae (formerly Sphingomonas terrae) is composed of 535 amino acid residues and one flavin adenine dinucleotide per monomer protein in a homodimeric structure. Its amino acid sequence shows 28.5 to 30.5% identity with glucose oxidases from Aspergillus niger and Penicillium amagasakiense. The ADP-binding site and the signature 1 and 2 consensus sequences of glucose-methanol-choline oxidoreductases are present in PEGDH. Based on three-dimensional molecular modeling and kinetic characterization of wild-type PEGDH and mutant PEGDHs constructed by site-directed mutagenesis, residues potentially involved in catalysis and substrate binding were found in the vicinity of the flavin ring. The catalytically important active sites were assigned to His-467 and Asn-511. One disulfide bridge between Cys-379 and Cys-382 existed in PEGDH and seemed to play roles in both substrate binding and electron mediation. The Cys-297 mutant showed decreased activity, suggesting the residue's importance in both substrate binding and electron mediation, as well as Cys-379 and Cys-382. PEGDH also contains a motif of a ubiquinone-binding site, and coenzyme Q10 was utilized as an electron acceptor. Thus, we propose several important amino acid residues involved in the electron transfer pathway from the substrate to ubiquinone.

Alcohol Oxidoreductases↗

Intrinsically disordered loops inserted into the structural domains of human proteins.

Much attention has been paid recently to proteins with partially or fully disordered structures, which are found to exist mostly in eukaryotes and are involved mainly in pivotal cellular processes such as transcriptional regulation, translation and cellular signal transduction. Long disordered sequences are sometimes inserted within the single structural domains of proteins, forming loops from the molecular surface. Such intrinsically disordered loops (IDLs) either are invisible in X-ray crystallography, or hamper protein crystallization itself due to great flexibility. Perhaps because of this, such long disordered sequences have not been characterized adequately. Here, we propose an informational method that stringently identifies IDLs in the structural domains of proteins using the amino acid sequence alone. A genome-wide survey of human proteins conducted with the method identified 50 IDL-containing proteins, several of which have experimentally determined 3D structures. Similar searches in other entirely sequenced organisms revealed that IDLs are prevalent in eukaryotes, while they are much less so in prokaryotes. As there is a statistically significant coincidence between the boundaries of IDLs and those of exons, we suggest that IDLs were produced mainly by exon addition in eukaryotes. IDLs are almost always located at the surface of proteins and are enriched with hydrophilic residues, and IDL-containing proteins tend to be intracellular. Some of the well-characterized proteins with IDLs illustrate that IDLs play pivotal roles in the switching of intracellular signaling or regulatory functions, suggesting that IDL insertion is an effective way to create functionally different domain variants.

Amino Acid Sequence↗

A complete set of Escherichia coli open reading frames in mobile plasmids facilitating genetic studies.

To facilitate genetic studies of Escherichia coli, we constructed a complete set of mobile plasmid clones of intact open reading frames (ORFs). Their expression is strictly controlled by Ptac / lacI(q). The plasmids carrying each ORF were introduced into an F+ recA strain and stored in 96-well microtiter plates. In this way, 96 clones can be transferred simultaneously to F- bacteria using the conjugative system. This provides a convenient procedure for systematic identification of ORFs that suppress or complement mutations. We created two types of clone sets: the original set contained individual clones in 45 microtiter plates, and a second set contained pools of 48 clones stored in a single microtiter plate. Using these clone sets, we have identified 403 genes that can correct in trans the temperature-sensitive defect of cell division mutants, which would suggest multiple global regulators for bacterial cell division.

Base Sequence↗

Recoverable one-dimensional encoding of three-dimensional protein structures.

One-dimensional (1D) structures of proteins such as secondary structure and contact number provide intuitive pictures to understand how the native three-dimensional (3D) structure of a protein is encoded in the amino acid sequence. However, it is still not clear whether a given set of 1D structures contains sufficient information for recovering the underlying 3D structure. Here we show that the 3D structure of a protein can be recovered from a set of three types of 1D structures, namely, secondary structure, contact number and residue-wise contact order which is introduced here for the first time. Using simulated annealing molecular dynamics simulations, the structures satisfying the given native 1D structural restraints were sought for 16 proteins of various structural classes and of sizes ranging from 56 to 146 residues. By selecting the structures best satisfying the restraints, all the proteins showed a coordinate RMS deviation of <4 A from the native structure, and, for most of them, the deviation was even <2 A. The present result opens a new possibility to protein structure prediction and our understanding of the sequence-structure relationship.

Amino Acid Sequence↗

Predicting absolute contact numbers of native protein structure from amino acid sequence.

The contact number of an amino acid residue in a protein structure is defined by the number of C(beta) atoms around the C(beta) atom of the given residue, a quantity similar to, but different from, solvent accessible surface area. We present a method to predict the contact numbers of a protein from its amino acid sequence. The method is based on a simple linear regression scheme and predicts the absolute values of contact numbers. When single sequences are used for both parameter estimation and cross-validation, the present method predicts the contact numbers with a correlation coefficient of 0.555 on average. When multiple sequence alignments are used, the correlation increases to 0.627, which is a significant improvement over previous methods. In terms of discrete states prediction, the accuracies for 2-, 3-, and 10-state predictions are, respectively, 71.4%, 54.1%, and 18.9% with residue type-dependent unbiased thresholds, and 76.3%, 59.2%, and 21.8% with residue type-independent unbiased thresholds. The difference between accessible surface area and contact number from a prediction viewpoint and the application of contact number prediction to three-dimensional structure prediction are discussed.

Amino Acid Sequence↗

Alternative splice variants encoding unstable protein domains exist in the human brain.

Alternative splicing has been recognized as a major mechanism by which protein diversity is increased without significantly increasing genome size in animals and has crucial medical implications, as many alternative splice variants are known to cause diseases. Despite the importance of knowing what structural changes alternative splicing introduces to the encoded proteins for the consideration of its significance, the problem has not been adequately explored. Therefore, we systematically examined the structures of the proteins encoded by the alternative splice variants in the HUGE protein database derived from long (>4 kb) human brain cDNAs. Limiting our analyses to reliable alternative splice junctions, we found alternative splice junctions to have a slight tendency to avoid the interior of SCOP domains and a strong statistically significant tendency to coincide with SCOP domain boundaries. These findings reflect the occurrence of some alternative splicing events that utilize protein structural units as a cassette. However, 50 cases were identified in which SCOP domains are disrupted in the middle by alternative splicing. In six of the cases, insertions are introduced at the molecular surface, presumably affecting protein functions, while in 11 of the cases alternatively spliced variants were found to encode pairs of stable and unstable proteins. The mRNAs encoding such unstable proteins are much less abundant than those encoding stable proteins and tend not to have corresponding mRNAs in non-primate species. We propose that most unstable proteins encoded by alternative splice variants lack normal functions and are an evolutionary dead-end.

Alternative Splicing↗

Estimation of the number of authentic orphan genes in bacterial genomes.

Genome annotation produces a considerable number of putative proteins lacking sequence similarity to known proteins. These are referred to as "orphans." The proportion of orphan genes varies among genomes, and is independent of genome size. In the present study, we show that the proportion of orphan genes roughly correlates with the isolation index of organisms (IIO), an indicator introduced in the present study, which represents the degree of isolation of a given genome as measured by sequence similarity. However, there are outlier genomes with respect to the linear correlation, consisting of those genomes that may contain excess amounts of orphan genes. Comparisons of genome sequences among closely related strains revealed that some of the annotated genes are not conserved, suggesting that they are ORFs occurring by chance. Exclusion of these non-conserved ORFs within closely related genomes improved the correlation between the proportion of orphan genes and the IIO values. Assuming that the correlation holds in general, this relationship was used to estimate the number of "authentic" orphan genes in a genome. Using this definition of authentic orphan genes, the anomalies arising from over-assignments, e.g., the percentages of structural annotations, were corrected for 16 genomes, including those of five archaea.

Amino Acid Sequence↗

Eigenvalue analysis of amino acid substitution matrices reveals a sharp transition of the mode of sequence conservation in proteins.

The pattern of amino acid substitutions and sequence conservation over many structure-based alignments of protein sequences was analyzed as a function of percentage sequence identity. The statistics of the amino acid substitutions were converted into the form of log-odds amino acid substitution matrices to which eigenvalue decomposition was applied. It was found that the most important component of the substitution matrices exhibited a sharp transition at the sequence identity of 30-35%, which coincides with the twilight zone. Above the transition point, the most dominant component is related to the mutability of amino acids and it acts to disfavor any substitutions, whereas below the transition point, the most dominant component is related to the hydrophobicity of amino acids and substitutions between residues of similar hydrophobic character are positively favored. Implications for protein evolution and sequence analysis are discussed.

Algorithms↗

Construction and characterization of chimeric proteins composed of type-1 and type-2 periplasmic binding proteins MglB and ArgT.

The respective type-1 and type-2 periplasmic binding proteins (PBPs) MglB and ArgT are believed to have evolved from a common ancestor into siblings showing topological differences in their main chain connectivity. At first glance, they show similar structure. But, more detailed examination reveals that the chain connectivity of ArgT is more convoluted than that of MglB. Reflecting that complexity, the folding of ArgT is complicated and involves intermediate folds. On the other hand, the folding of MglB is a simple two-state transition. In the present study, we constructed and characterized several chimeras made up of various subdomains of MglB and ArgT with the aim of gaining insight into the evolution of protein folding and protein structure. Although these chimeras did not fold as compactly as their parental proteins, some did exhibit cooperative folding, which suggests that novel proteins with new connectivity and new folding pathways could have emerged at a fairly high rate throughout the evolution of proteins.

Circular Dichroism↗

Prediction of catalytic residues in enzymes based on known tertiary structure, stability profile, and sequence conservation.

The catalytic or functionally important residues of a protein are known to exist in evolutionarily constrained regions. However, the patterns of residue conservation alone are sometimes not very informative, depending on the homologous sequences available for a given query protein. Here, we present an integrated method to locate the catalytic residues in an enzyme from its sequence and structure. Mutations of functional residues usually decrease the activity, but concurrently often increase stability. Also, catalytic residues tend to occupy partially buried sites in holes or clefts on the molecular surface. After confirming these general tendencies by carrying out statistical analyses on 49 representative enzymes, these data together with amino acid conservation were evaluated. This novel method exhibited better sensitivity in the prediction accuracy than traditional methods that consider only the residue conservation. We applied it to some so-called "hypothetical" proteins, with known structures but undefined functions. The relationships among the catalytic, conserved, and destabilizing residues in enzymatic proteins are discussed.

Amino Acid Sequence↗

Unique amino acid composition of proteins in halophilic bacteria.

The amino acid compositions of proteins from halophilic archaea were compared with those from non-halophilic mesophiles and thermophiles, in terms of the protein surface and interior, on a genome-wide scale. As we previously reported for proteins from thermophiles, a biased amino acid composition also exists in halophiles, in which an abundance of acidic residues was found on the protein surface as compared to the interior. This general feature did not seem to depend on the individual protein structures, but was applicable to all proteins encoded within the entire genome. Unique protein surface compositions are common in both halophiles and thermophiles. Statistical tests have shown that significant surface compositional differences exist among halophiles, non-halophiles, and thermophiles, while the interior composition within each of the three types of organisms does not significantly differ. Although thermophilic proteins have an almost equal abundance of both acidic and basic residues, a large excess of acidic residues in halophilic proteins seems to be compensated by fewer basic residues. Aspartic acid, lysine, asparagine, alanine, and threonine significantly contributed to the compositional differences of halophiles from meso- and thermophiles. Among them, however, only aspartic acid deviated largely from the expected amount estimated from the dinucleotide composition of the genomic DNA sequence of the halophile, which has an extremely high G+C content (68%). Thus, the other residues with large deviations (Lys, Ala, etc.) from their non-halophilic frequencies could have arisen merely as "dragging effects" caused by the compositional shift of the DNA, which would have changed to increase principally the fraction of aspartic acid alone.

Amino Acids↗