Search PubMed⌕ Search

Biomedical subjects

Daniel Falush

Publications and source records attributed to Daniel Falush.

13 recordsLinked to original sources

Inference of bacterial microevolution using multilocus sequence data.

We describe a model-based method for using multilocus sequence data to infer the clonal relationships of bacteria and the chromosomal position of homologous recombination events that disrupt a clonal pattern of inheritance. The key assumption of our model is that recombination events introduce a constant rate of substitutions to a contiguous region of sequence. The method is applicable both to multilocus sequence typing (MLST) data from a few loci and to alignments of multiple bacterial genomes. It can be used to decide whether a subset of isolates share common ancestry, to estimate the age of the common ancestor, and hence to address a variety of epidemiological and ecological questions that hinge on the pattern of bacterial spread. It should also be useful in associating particular genetic events with the changes in phenotype that they cause. We show that the model outperforms existing methods of subdividing recombinogenic bacteria using MLST data and provide examples from Salmonella and Bacillus. The software used in this article, ClonalFrame, is available from http://bacteria.stats.ox.ac.uk/.

Bacteria↗

Mismatch induced speciation in Salmonella: model and data.

In bacteria, DNA sequence mismatches act as a barrier to recombination between distantly related organisms and can potentially promote the cohesion of species. We have performed computer simulations which show that the homology dependence of recombination can cause de novo speciation in a neutrally evolving population once a critical population size has been exceeded. Our model can explain the patterns of divergence and genetic exchange observed in the genus Salmonella, without invoking either natural selection or geographical population subdivision. If this model was validated, based on extensive sequence data, it would imply that the named subspecies of Salmonella enterica correspond to good biological species, making species boundaries objective. However, multilocus sequence typing data, analysed using several conventional tools, provide a misleading impression of relationships within S. enterica subspecies enterica and do not provide the resolution to establish whether new species are presently being formed.

Base Pair Mismatch↗

A bimodal pattern of relatedness between the Salmonella Paratyphi A and Typhi genomes: convergence or divergence by homologous recombination?

All Salmonella can cause disease but severe systemic infections are primarily caused by a few lineages. Paratyphi A and Typhi are the deadliest human restricted serovars, responsible for approximately 600,000 deaths per annum. We developed a Bayesian changepoint model that uses variation in the degree of nucleotide divergence along two genomes to detect homologous recombination between these strains, and with other lineages of Salmonella enterica. Paratyphi A and Typhi showed an atypical and surprising pattern. For three quarters of their genomes, they appear to be distantly related members of the species S. enterica, both in their gene content and nucleotide divergence. However, the remaining quarter is much more similar in both aspects, with average nucleotide divergence of 0.18% instead of 1.2%. We describe two different scenarios that could have led to this pattern, convergence and divergence, and conclude that the former is more likely based on a variety of criteria. The convergence scenario implies that, although Paratyphi A and Typhi were not especially close relatives within S. enterica, they have gone through a burst of recombination involving more than 100 recombination events. Several of the recombination events transferred novel genes in addition to homologous sequences, resulting in similar gene content in the two lineages. We propose that recombination between Typhi and Paratyphi A has allowed the exchange of gene variants that are important for their adaptation to their common ecological niche, the human host.

Algorithms↗

Genome-wide association mapping in bacteria?

Bacteria display many interesting phenotypes such as virulence, tissue specificity and host range, for which it would be useful to know the genetic basis. Association mapping involves identifying causal variants by showing that particular genotypes are statistically associated with a phenotypic trait in a sample of strains taken from a natural population. With the advent of high-throughput genotyping, association mapping is becoming an increasingly powerful approach. However, until recently, association studies had not been used in bacteria because of their strong population structure, which can produce false positives and/or loss of statistical power unless elucidated and taken into account in analyses. Here, we describe how association mapping could be successfully applied to bacteria and outline the necessary sampling and genotyping strategies.

Bacteria↗

Sex and virulence in Escherichia coli: an evolutionary perspective.

Pathogenic Escherichia coli cause over 160 million cases of dysentery and one million deaths per year, whereas non-pathogenic E. coli constitute part of the normal intestinal flora of healthy mammals and birds. The evolutionary pathways underlying this dichotomy in bacterial lifestyle were investigated by multilocus sequence typing of a global collection of isolates. Specific pathogen types [enterohaemorrhagic E. coli, enteropathogenic E. coli, enteroinvasive E. coli, K1 and Shigella] have arisen independently and repeatedly in several lineages, whereas other lineages contain only few pathogens. Rates of evolution have accelerated in pathogenic lineages, culminating in highly virulent organisms whose genomic contents are altered frequently by increased rates of homologous recombination; thus, the evolution of virulence is linked to bacterial sex. This long-term pattern of evolution was observed in genes distributed throughout the genome, and thereby is the likely result of episodic selection for strains that can escape the host immune response.

Alleles↗

Genomic changes during chronic Helicobacter pylori infection.

The gastric pathogen Helicobacter pylori shows tremendous genetic variability within human populations, both in gene content and at the sequence level. We investigated how this variability arises by comparing the genome content of 21 closely related pairs of isolates taken from the same patient at different time points. The comparisons were performed by hybridization with whole-genome DNA microarrays. All loci where microarrays indicated a genomic change were sequenced to confirm the events. The number of genomic changes was compared to the number of homologous replacement events without loss or gain of genes that we had previously determined by multilocus sequence analysis and mathematical modeling based on the sequence data. Our analysis showed that the great majority of genetic changes were due to homologous recombination, with 1/650 events leading to a net gain or loss of genes. These results suggest that adaptation of H. pylori to the host individual may principally occur through sequence changes rather than loss or gain of genes.

Adult↗

Sequence typing and comparison of population biology of Campylobacter coli and Campylobacter jejuni.

A multilocus sequence typing (MLST) scheme that uses the same loci as a previously described system for Campylobacter jejuni was developed for Campylobacter coli. The C. coli-specific primers were validated with 53 isolates from humans, chickens, and pigs, together with 15 Penner serotype reference isolates. The nucleotide sequence of the flaA short variable region (SVR) was determined for each isolate. These sequence data were compared to equivalent information for 17 C. jejuni isolates representing the known genetic diversity of this species. C. coli and C. jejuni share approximately 86.5% identity at the nucleotide sequence level within the MLST loci. There is evidence of genetic exchange of the housekeeping genes between the two species, but at a very low rate; only one sequence type from each species showed evidence of imported DNA. The flaA gene was more variable and has been exchanged many times between the two species, making it an unreliable marker for species identification but useful for distinguishing closely related strains. All but 3 of 21 human C. coli clinical isolates were distinct, according to the combined MLST and SVR sequences. The use of a common MLST scheme allows direct comparisons of the population biology and molecular epidemiology of these two closely related human pathogens.

Animals↗

Hybrid Vibrio vulnificus.

The recent emergence of the human-pathogenic Vibrio vulnificus in Israel was investigated by using multilocus genotype data and modern molecular evolutionary analysis tools. We show that this pathogen is a hybrid organism that evolved by the hybridization of the genomes from 2 distinct and independent populations. These findings provide clear evidence of how hybridization between 2 existing and nonpathogenic forms has apparently led to the emergence of an epidemic infectious disease caused by this pathogenic variant. This novel observation shows yet another way in which epidemic organisms arise.

Animals↗

Germs, genomes and genealogies.

Genetic diversity in pathogen species contains information about evolutionary and epidemiological processes, including the origins and history of disease, the nature of the selective forces acting on pathogen genes and the role of recombination in generating genetic novelty. Here, we review recent developments in these fields and compare the use of population genetic, or population-model based, approaches to phylogenetic, or population-model free, methodologies. We show how simple epidemiological models can be related to the ancestral, or coalescent, process underlying samples from pathogen species, enabling detailed inference about pathogen biology from patterns of molecular variation.

Journal Article↗

Origin of extant domesticated sunflowers in eastern North America.

Eastern North America is one of at least six regions of the world where agriculture is thought to have arisen wholly independently. The primary evidence for this hypothesis derives from morphological changes in the archaeobotanical record of three important crops--squash, goosefoot and sunflower--as well as an extinct minor cultigen, sumpweed. However, the geographical origins of two of the three primary domesticates--squash and goosefoot--are now debated, and until recently sunflower (Helianthus annuus L.) has been considered the only undisputed eastern North American domesticate. The discovery of 4,000-year-old domesticated sunflower remains from San Andrés, Tabasco, implies an earlier and possibly independent origin of domestication in Mexico and has stimulated a re-examination of the geographical origin of domesticated sunflower. Here we describe the genetic relationships and pattern of genetic drift between extant domesticated strains and wild populations collected from throughout the USA and Mexico. We show that extant domesticates arose in eastern North America, with a substantial genetic bottleneck occurring during domestication.

Agriculture↗

Distinguishing human ethnic groups by means of sequences from Helicobacter pylori: lessons from Ladakh.

The history of mankind remains one of the most challenging fields of study. However, the emergence of anatomically modern humans has been so recent that only a few genetically informative polymorphisms have accumulated. Here, we show that DNA sequences from Helicobacter pylori, a bacterium that colonizes the stomachs of most humans and is usually transmitted within families, can distinguish between closely related human populations and are superior in this respect to classical human genetic markers. H. pylori from Buddhists and Muslims, the two major ethnic communities in Ladakh (India), differ in their population-genetic structure. Moreover, the prokaryotic diversity is consistent with the Buddhists having arisen from an introgression of Tibetan speakers into an ancient Ladakhi population. H. pylori from Muslims contain a much stronger ancestral Ladakhi component, except for several isolates with an Indo-European signature, probably reflecting genetic flux from the Near East. These signatures in H. pylori sequences are congruent with the recent history of population movements in Ladakh, whereas similar signatures in human microsatellites or mtDNA were only marginally significant. H. pylori sequence analysis has the potential to become an important tool for unraveling short-term genetic changes in human populations.

DNA, Bacterial↗

Traces of human migrations in Helicobacter pylori populations.

Helicobacter pylori, a chronic gastric pathogen of human beings, can be divided into seven populations and subpopulations with distinct geographical distributions. These modern populations derive their gene pools from ancestral populations that arose in Africa, Central Asia, and East Asia. Subsequent spread can be attributed to human migratory fluxes such as the prehistoric colonization of Polynesia and the Americas, the neolithic introduction of farming to Europe, the Bantu expansion within Africa, and the slave trade.

Africa↗

Inference of population structure using multilocus genotype data: linked loci and correlated allele frequencies.

We describe extensions to the method of Pritchard et al. for inferring population structure from multilocus genotype data. Most importantly, we develop methods that allow for linkage between loci. The new model accounts for the correlations between linked loci that arise in admixed populations ("admixture linkage disequilibium"). This modification has several advantages, allowing (1) detection of admixture events farther back into the past, (2) inference of the population of origin of chromosomal regions, and (3) more accurate estimates of statistical uncertainty when linked loci are used. It is also of potential use for admixture mapping. In addition, we describe a new prior model for the allele frequencies within each population, which allows identification of subtle population subdivisions that were not detectable using the existing method. We present results applying the new methods to study admixture in African-Americans, recombination in Helicobacter pylori, and drift in populations of Drosophila melanogaster. The methods are implemented in a program, structure, version 2.0, which is available at http://pritch.bsd.uchicago.edu.

Algorithms↗