Search PubMed⌕ Search

Biomedical subjects

Hirohisa Kishino

Publications and source records attributed to Hirohisa Kishino.

At least 19 recordsLinked to original sources

Phylogenetic methodology for detecting protein interactions.

Detecting protein-protein interactions and assigning proteins to functional complexes are key challenges of modern biology. The rise of genomics has lead to evidence that correlated patterns of presence/absence and/or fusing of proteins in any organism suggest these proteins interact. Unfortunately, methods based on such data work best with divergent genomes, whereas major sequencing efforts in vertebrates, for example, are yielding alignments of the same set of proteins sampled from the same set of taxa (species). Using vertebrate mitochondrial genomes to illustrate a novel method, we associate proteins based on vectors of their evolutionary tree edge (branch or internode) lengths. This approach is based on the expectation that molecular coevolution is greatest between proteins that interact in some way. Mitochondrial DNA-encoded proteins are associated into groups largely consistent with the complexes they come from. This association is apparently not due to the tree structure or mutation processes, leaving coevolution as the best explanation. We show that it is important that the tree used to derive the edge-length vector is estimated accurately in terms of both topology and edge lengths. Although more complex substitution models reduce systematic error, they also inflate stochastic error. This makes the use of less complex substitution models preferable in some circumstances. We describe a method to estimate correlations of pairwise evolutionary distances, which adjusts for non-independent correlations due to shared evolutionary history. Associations of proteins based on their edge-length vectors are visualized and assessed using a variety of hierarchical clustering and multidimensional scaling methods. New formula for estimating the fit of data to model, including the average percent standard deviation of distances on least squares trees, are presented. Use of edge-length vectors is compared and contrasted with correlated distance methods, correlated rates methods, and site-specific evidence of coevolution.

Cluster Analysis↗

An integrated-likelihood method for estimating genetic differentiation between populations.

The aim of this article is to develop an integrated-likelihood (IL) approach to estimate the genetic differentiation between populations. The conventional maximum-likelihood (ML) and pseudolikelihood (PL) methods that use sample counts of alleles may cause severe underestimations of FST, which means overestimations of theta=4Nm, when the number of sampling localities is small. To reduce such bias in the estimation of genetic differentiation, we propose an IL method in which the mean allele frequencies over populations are regarded as nuisance parameters and are eliminated by integration. To maximize the IL function, we have developed two algorithms, a Monte Carlo EM algorithm and a Laplace approximation. Our simulation studies show that the method proposed here outperforms the conventional ML and PL methods in terms of unbiasedness and precision. The IL method was applied to real data for Pacific herring and African elephants.

Algorithms↗

Core set approach to reduce uncertainty of gene trees.

BACKGROUND: A genealogy based on gene sequences within a species plays an essential role in the estimation of the character, structure, and evolutionary history of that species. Because intraspecific sequences are more closely related than interspecific ones, detailed information on the evolutionary process may be available by determining all the node sequences of trees and provide insight into functional constraints and adaptations. However, strong evolutionary correlations on a few lineages make this determination difficult as a whole, and the maximum parsimony (MP) method frequently allows a number of topologies with a same total branching length. RESULTS: Kitazoe et al. developed multidimensional vector-space representation of phylogeny. It converts additivity of evolutionary distances to orthogonality among the vectors expressing branches, and provides a unified index to measure deviations from the orthogoality. In this paper, this index is used to detect and exclude sequences with large deviations from orthogonality, and then selects a maximum subset ("core set") of sequences for which MP generates a single solution. Once the core set tree is formed whose all the node sequences are given, the excluded sequences are found to have basically two phylogenetic positions on this tree, respectively. Fortunately, since multiple substitutions are rare in intra-species sequences, the variance of nucleotide transitions is confined to a small range. By applying the core set approach to 38 partial env sequences of HIV-1 in a single patient and also 198 mitochondrial COI and COII DNA sequences of Anopheles dirus, we demonstrate how consistently this approach constructs the tree. CONCLUSION: In the HIV dataset, we confirmed that the obtained core set tree is the unique maximum set for which MP proposes a single tree. In the mosquito data set, the fluctuation of nucleotide transitions caused by the sequences excluded from the core set was very small. We reproduced this core-set tree by simulation based on random process, and applied our approach to many sets of the obtained endpoint sequences. Consequently, the ninety percent of the endpoint sequences was identified as the core sets and the obtained node sequences were perfectly identical to the true ones.

Animals↗

Positive selection acting on a surface membrane protein of the plant-pathogenic phytoplasmas.

Phytoplasmas are plant-pathogenic bacteria that cause numerous diseases. This study shows a strong positive selection on the phytoplasma antigenic membrane protein (Amp). The ratio of nonsynonymous to synonymous substitutions was >1 with all the methods we tested. The clear positive selections imply an important biological role for Amp in host-bacterium interactions.

Antigens, Bacterial↗

Simultaneous estimation of mixing rates and genetic drift under successive sampling of genetic markers with application to the mud crab (Scylla paramamosain) in Japan.

In stock enhancement programs, it is important to assess mixing rates of released individuals in stocks. For this purpose, genetic stock identification has been applied. The allele frequencies in a composite population are expressed as a mixture of the allele frequencies in the natural and released populations. The estimation of mixing rates is possible, under successive sampling from the composite population, on the basis of temporal changes in allele frequencies. The allele frequencies in the natural population may be estimated from those of the composite population in the preceding year. However, it should be noted that these frequencies can vary between generations due to genetic drift. In this article, we develop a new method for simultaneous estimation of mixing rates and genetic drift in a stock enhancement program. Numerical simulation shows that our procedure estimates the mixing rate with little bias. Although the genetic drift is underestimated when the amount of information is small, reduction of the bias is possible by analyzing multiple unlinked loci. The method was applied to real data on mud crab stocking, and the result showed a yearly variation in the mixing rate.

Animals↗

Fold recognition of the human immunodeficiency virus type 1 V3 loop and flexibility of its crown structure during the course of adaptation to a host.

The third hypervariable (V3) region of the HIV-1 gp120 protein is responsible for many aspects of viral infectivity. The tertiary structure of the V3 loop seems to influence the coreceptor usage of the virus, which is an important determinant of HIV pathogenesis. Hence, the information about preferred conformations of the V3-loop region and its flexibility could be a crucial tool for understanding the mechanisms of progression from an initial infection to AIDS. Taking into account the uncertainty of the loop structure, we predicted the structural flexibility, diversity, and sequence fitness to the V3-loop structure for each of the sequences serially sampled during an asymptomatic period. Structural diversity correlated with sequence diversity. The predicted crown structure usage implied that structural flexibility depended on the patient and that the antigenic character of the virus might be almost uniform in a patient whose immune system is strong. Furthermore, the predicted structural ensemble suggested that toward the end of the asymptomatic period there was a change in the V3-loop structure or in the environment surrounding the V3 loop, possibly because of its proximity to the gp120 core.

Adaptation, Physiological↗

Incorporating gene-specific variation when inferring and evaluating optimal evolutionary tree topologies from multilocus sequence data.

Because of the increase of genomic data, multiple genes are often available for the inference of phylogenetic relationships. The simple approach for combining multiple genes from the same taxon is to concatenate the sequences and then ignore the fact that different positions in the concatenated sequence came from different genes. Here, we discuss two criteria for inferring the optimal tree topology from data sets with multiple genes. These criteria are designed for multigene data sets where gene-specific evolutionary features are too important to ignore. One criterion is conventional and is obtained by taking the sum of log-likelihoods over all genes. The other criterion is obtained by dividing the log-likelihood for a gene by its sequence length and then taking the arithmetic mean over genes of these ratios. A similar strategy could be adopted with parsimony scores. The optimal tree is then declared to be the one for which the sum or the arithmetic mean is maximized. These criteria are justified within a two-stage hierarchical framework. The first level of the hierarchy represents gene-specific evolutionary features, and the second represents site-specific features for given genes. For testing significance of the optimal topology, we suggest a two-stage bootstrap procedure that involves resampling genes and then resampling alignment columns within resampled genes. An advantage of this procedure over concatenation is that it can effectively account for gene-specific evolutionary features. We discuss the applicability of the two-stage bootstrap idea to the Kishino-Hasegawa test and the Shimodaira-Hasegawa test.

Biological Evolution↗

Multidimensional vector space representation for convergent evolution and molecular phylogeny.

With growing amounts of genome data and constant improvement of models of molecular evolution, phylogenetic reconstruction became more reliable. However, our knowledge of the real process of molecular evolution is still limited. When enough large-sized data sets are analyzed, any subtle biases in statistical models can support incorrect topologies significantly because of the high signal-to-noise ratio. We propose a procedure to locate sequences in a multidimensional vector space (MVS), in which the geometry of the space is uniquely determined in such a way that the vectors of sequence evolution are orthogonal among different branches. In this paper, the MVS approach is developed to detect and remove biases in models of molecular evolution caused by unrecognized convergent evolution among lineages or unexpected patterns of substitutions. Biases in the estimated pairwise distances are identified as deviations (outliers) of sequence spatial vectors from the expected orthogonality. Modifications to the estimated distances are made by minimizing an index to quantify the deviations. In this way, it becomes possible to reconstruct the phylogenetic tree, taking account of possible biases in the model of molecular evolution. The efficacy of the modification procedure was verified by simulating evolution on various topologies with rate heterogeneity and convergent change. The phylogeny of placental mammals in previous analyses of large data sets has varied according to the genes being analyzed. Systematic deviations caused by convergent evolution were detected by our procedure in all representative data sets and were found to strongly affect the tree structure. However, the bias correction yielded a consistent topology among data sets. The existence of strong biases was validated by examining the sites of convergent evolution between the hedgehog and other species in mitochondrial data set. This convergent evolution explains why it has been difficult to determine the phylogenetic placement of the hedgehog in previous studies.

Animals↗

Divergence pattern of duplicate genes in protein-protein interactions follows the power law.

The impact of the biological network structures on the divergence between the two copies of one duplicate gene pair involved in the networks has not been documented on a genome scale. Having analyzed the most recently updated Database of Interacting Proteins (DIP) by incorporating the information for duplicate genes of the same age in yeast, we find that there was a highly significantly positive correlation between the level of connectivity of ancient genes and the number of shared partners of their duplicates in the protein-protein interaction networks. This suggests that duplicate genes with a low ancestral connectivity tend to provide raw materials for functional novelty, whereas those duplicate genes with a high ancestral connectivity tend to create functional redundancy for a genome during the same evolutionary period. Moreover, the difference in the number of partners between two copies of a duplicate pair was found to follow a power-law distribution. This suggests that loss and gain of interacting partners for most duplicate genes with a lower level of ancestral connectivity is largely symmetrical, whereas the "hub duplicate genes" with a higher level of ancient connectivity display an asymmetrical divergence pattern in protein-protein interactions. Thus, it is clear that the protein-protein interaction network structures affect the divergence pattern of duplicate genes. Our findings also provide insights into the origin and development of biological networks.

Databases, Protein↗

Estimating absolute rates of synonymous and nonsynonymous nucleotide substitution in order to characterize natural selection and date species divergences.

The rate of molecular evolution can vary among lineages. Sources of this variation have differential effects on synonymous and nonsynonymous substitution rates. Changes in effective population size or patterns of natural selection will mainly alter nonsynonymous substitution rates. Changes in generation length or mutation rates are likely to have an impact on both synonymous and nonsynonymous substitution rates. By comparing changes in synonymous and nonsynonymous rates, the relative contributions of the driving forces of evolution can be better characterized. Here, we introduce a procedure for estimating the chronological rates of synonymous and nonsynonymous substitutions on the branches of an evolutionary tree. Because the widely used ratio of nonsynonymous and synonymous rates is not designed to detect simultaneous increases or simultaneous decreases in synonymous and nonsynonymous rates, the estimation of these rates rather than their ratio can improve characterization of the evolutionary process. With our Bayesian approach, we analyze cytochrome oxidase subunit I evolution in primates and infer that nonsynonymous rates have a greater tendency to change over time than do synonymous rates. Our analysis of these data also suggests that rates have been positively correlated.

Animals↗

Simultaneous detection of linkage disequilibrium and genetic differentiation of subdivided populations.

We propose a new method for simultaneously detecting linkage disequilibrium and genetic structure in subdivided populations. Taking subpopulation structure into account with a hierarchical model, we estimate the magnitude of genetic differentiation and linkage disequilibrium in a metapopulation on the basis of geographical samples, rather than decompose a population into a finite number of random-mating subpopulations. We assume that Hardy-Weinberg equilibrium is satisfied in each locality, but do not assume independence between marker loci. Linkage states remain unknown. Genetic differentiation and linkage disequilibrium are expressed as hyperparameters describing the prior distribution of genotypes or haplotypes. We estimate related parameters by maximizing marginal-likelihood functions and detect linkage equilibrium or disequilibrium by the Akaike information criterion. Our empirical Bayesian model analyzes genotype and haplotype frequencies regardless of haploid or diploid data, so it can be applied to most commonly used genetic markers. The performance of our procedure is examined via numerical simulations in comparison with classical procedures. Finally, we analyze isozyme data of ayu, a severely exploited fish species, and single-nucleotide polymorphisms in human ALDH2.

Animals↗

Genomic background predicts the fate of duplicated genes: evidence from the yeast genome.

Gene duplication with subsequent divergence plays a central role in the acquisition of genes with novel function and complexity during the course of evolution. With reduced functional constraints or through positive selection, these duplicated genes may experience accelerated evolution. Under the model of subfunctionalization, loss of subfunctions leads to complementary acceleration at sites with two copies, and the difference in average rate between the sequences may not be obvious. On the other hand, the classical model of neofunctionalization predicts that the evolutionary rate in one of the two duplicates is accelerated. However, the classical model does not tell which of the duplicates experiences the acceleration in evolutionary rate. Here, we present evidence from the Saccharomyces cerevisiae genome that a duplicate located in a genomic region with a low-recombination rate is likely to evolve faster than a duplicate in an area of high recombination. This observation is consistent with population genetics theory that predicts that purifying selection is less effective in genomic regions of low recombination (Hill-Robertson effect). Together with previous studies, our results suggest the genomic background (e.g., local recombination rate) as a potential force to drive the divergence between nontandemly duplicated genes. This implies the importance of structure and complexity of genomes in the diversification of organisms via gene duplications.

Evolution, Molecular↗

Genomic background drives the divergence of duplicated amylase genes at synonymous sites in Drosophila.

In some Drosophila species, there are two types of greatly diverged amylase (Amy) genes (Amy clusters 1 and 2), each encoding active amylase isozymes. Cluster 1 is located at the middle of its chromosomal arm, and the region has a normal local recombination rate. However, cluster 2 is near the centromere, and this region is known to have a reduced recombination rate. Although nonsynonymous substitutions follow a molecular clock, synonymous substitutions were accelerated in cluster 2 after gene duplications. This resulted in a higher GC content at the third codon position (GC3) and codon usage bias in cluster 1, and lower GC3 content and codon usage bias in the cluster 2. However, no systematic difference in GC content was observed in the first and second codon positions or the 3'-flanking regions. Therefore, differences in local recombination rate rather than mutation bias might explain the divergence at synonymous sites between the two Amy clusters within species (Hill-Robertson effect). Alternatively, the different patterns and levels of expression between the two clusters may imply that the reduced expression level in cluster 2 caused by chromatin potentiation decreased the codon bias. Both of these hypotheses imply the importance of the genomic background as a driving force of divergence between non-tandemly duplicated genes.

3' Flanking Region↗

Protein evolution with dependence among codons due to tertiary structure.

Markovian models of protein evolution that relax the assumption of independent change among codons are considered. With this comparatively realistic framework, an evolutionary rate at a site can depend both on the state of the site and on the states of surrounding sites. By allowing a relatively general dependence structure among sites, models of evolution can reflect attributes of tertiary structure. To quantify the impact of protein structure on protein evolution, we analyze protein-coding DNA sequence pairs with an evolutionary model that incorporates effects of solvent accessibility and pairwise interactions among amino acid residues. By explicitly considering the relationship between nonsynonymous substitution rates and protein structure, this approach can lead to refined detection and characterization of positive selection. Analyses of simulated sequence pairs indicate that parameters in this evolutionary model can be well estimated. Analyses of lysozyme c and annexin V sequence pairs yield the biologically reasonable result that amino acid replacement rates are higher when the replacements lead to energetically favorable proteins than when they destabilize the proteins. Although the focus here is evolutionary dependence among codons that is associated with protein structure, the statistical approach is quite general and could be applied to diverse cases of evolutionary dependence where surrogates for sequence fitness can be measured or modeled.

Annexin A5↗

Evolutionary history and mode of the amylase multigene family in Drosophila.

Previous studies indicate that the tandemly repeated members of the amylase (Amy) gene family evolved in a concerted manner in the melanogaster subgroup and in some other species. In this paper, we analyzed all of the 49 active and complete Amy gene sequences in Drosophila, mostly from subgenus Sophophora. Phylogenetic analysis indicated that the two types of diverged Amy genes in the Drosophila montium subgroup and Drosophila ananassae, which are located in distant chromosomal regions from each other, originated independently in different evolutionary lineages of the melanogaster group after the split of the obscura and melanogaster groups. One of the two clusters was lost after duplication in the melanogaster subgroup. Given the time, 24.9 mya, of divergence between the obscura and the melanogaster groups (Russo et al. 1995), the two duplication events were estimated to occur at about 13.96 +/- 1.93 and 12.38 +/- 1.76 mya in the montium subgroup and D. ananassae, respectively. An accelerated rate of amino acid changes was not observed in either lineage after these gene duplications. However, the G+C contents at the third codon positions (GC3) decreased significantly along one of the two Amy clusters both in the montium subgroup and in D. ananassae right after gene duplication. Furthermore, one of the two types of the Amy genes with a lower GC3 content has lost a specific regulatory element within the montium subgroup species and D. ananassae. While the tandemly repeated members evolved in a concerted manner, the two types of diverged Amy genes in Drosophila experienced frequent gene duplication, gene loss, and divergent evolution following the model of a birth-and-death process.

Amylases↗

Partial conservation of LFY function between rice and Arabidopsis.

The LFY/FLO genes encode plant-specific transcription factors and play major roles in the reproductive transition as well as floral development. In this study, we reconstructed the phylogenetic tree of the 49 LFY/FLO homologs from various plant species. The tree clearly shows that the LFY/FLO genes from the eudicots and monocots formed the two monophyletic clusters with very high bootstrap probabilities, respectively. Furthermore, grass LFY/FLO genes have experienced significant acceleration of amino acid replacement rate compared with the eudicot homolog. To test whether grass LFY/FLO genes have a conserved function with those of eudicots, we introduced RFL, a rice LFY homolog, into the Arabidopsis lfy mutant. The RFL gene driven by LFY promoter partially rescued the lfy mutation, suggesting that the functions of LFY and RFL partly overlap. Interestingly, the RFL but not LFY, strongly activated the expression of AP1 and AG, the downstream targets of LFY, even in the vegetative tissues. The LFY::RFL transgenic Arabidopsis plants exhibited abnormal patterns of development such as leaf curling, bushy appearance and the transformation of ovules into carpels. All of the results indicate that both the partial conservation and divergence of LFY function between rice and Arabidopsis.

Amino Acid Sequence↗

Time scale of eutherian evolution estimated without assuming a constant rate of molecular evolution.

Controversies over the molecular clock hypothesis were reviewed. Since it is evident that the molecular clock does not hold in an exact sense, accounting for evolution of the rate of molecular evolution is a prerequisite when estimating divergence times with molecular sequences. Recently proposed statistical methods that account for this rate variation are overviewed and one of these procedures is applied to the mitochondrial protein sequences and to the nuclear gene sequences from many mammalian species in order to estimate the time scale of eutherian evolution. This Bayesian method not only takes account of the variation of molecular evolutionary rate among lineages and among genes, but it also incorporates fossil evidence via constraints on node times. With denser taxonomic sampling and a more realistic model of molecular evolution, this Bayesian approach is expected to increase the accuracy of divergence time estimates.

Animals↗

Time flies, a new molecular time-scale for brachyceran fly evolution without a clock.

The insect order Diptera, the true flies, contains one of the four largest Mesozoic insect radiations within its suborder Brachycera. Estimates of phylogenetic relationships and divergence dates among the major brachyceran lineages have been problematic or vague because of a lack of consistent evidence and the rarity of well-preserved fossils. Here, we combine new evidence from nucleotide sequence data, morphological reinterpretations, and fossils to improve estimates of brachyceran evolutionary relationships and ages. The 28S ribosomal DNA (rDNA) gene was sequenced for a broad diversity of taxa, and the data were combined with recently published morphological scorings for a parsimony-based phylogenetic analysis. The phylogenetic topology inferred from the combined 28S rDNA and morphology data set supports brachyceran monophyly and the monophyly of the four major brachyceran infraorders and suggests relationships largely consistent with previous classifications. Weak support was found for a basal brachyceran clade comprising the infraorders Stratiomyomorpha (soldier flies and relatives), Xylophagomorpha (xylophagid flies), and Tabanomorpha (horse flies, snipe flies, and relatives). This topology and similar alternative arrangements were used to obtain Bayesian estimates of divergence times, both with and without the assumption of a constant evolutionary rate. The estimated times were relatively robust to the choice of prior distributions. Divergence times based on the 28S rDNA and several fossil constraints indicate that the Brachycera originated in the late Triassic or earliest Mesozoic and that all major lower brachyceran fly lineages had near contemporaneous origins in the mid-Jurassic prior to the origin of flowering plants (angiosperms). This study provides increased resolution of brachyceran phylogeny, and our revised estimates of fly ages should improve the temporal context of evolutionary inferences and genomic comparisons between fly model organisms.

Animals↗