Search PubMedSearch

SEARCH · Search PubMed

Results for “Lineage tree”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

LAML-Pro: joint maximum likelihood inference of cell genotypes and cell lineage trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (i) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (ii) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈25%-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is implemented in C++ and is available as both a command-line interface and as a Python library at: github.com/raphael-group/LAML-Pro.

Cell Lineage

LAML-Pro: Joint Maximum Likelihood Inference of Cell Genotypes and Cell Lineage Trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (1) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (2) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈ 25-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is freely available at: github.com/raphael-group/LAML-Pro.

Journal Article

Bayesian inference of lineage trees by joint analysis of single-cell multimodal lineage-tracing data with BiLinT.

The advent of single-cell lineage-tracing technologies has enabled the simultaneous profiling of gene expression and lineage barcodes. However, accurate, high-resolution reconstruction of cell lineage trees remains challenging because most existing approaches treat these modalities separately and therefore fail to fully exploit their complementary information. Here we present BiLinT, a Bayesian framework that jointly models multimodal single-cell lineage-tracing data for lineage tree reconstruction. BiLinT integrates barcode evolution (a continuous-time Markov chain) with gene expression dynamics (an Ornstein-Uhlenbeck process) within a unified probabilistic model. Across synthetic and real data sets, BiLinT provides accurate lineage-tree reconstruction and reveals differentiation-associated clonal structure and developmental fate biases.

Journal Article

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis

Tracking Somatic Mutations for Lineage Reconstruction.

The human genome is composed of distinct genomic regions that are susceptible to various types of somatic mutations. Among these, Short Tandem Repeats (STRs) stand out as the most mutable genetic elements. STRs are short repetitive polymorphic sequences, predominantly situated within noncoding sectors of the genome. The intrinsic repetition characterizing these sequences makes them highly mutable in vivo. Consequently, this characteristic provides the chance to unravel the natural developmental history of human viable cells retrospectively. However, STRs also introduce stutter noise in vitro amplification, which makes their analysis challenging. Here we describe our integrated biochemical-computational platform for single-cell lineage analysis. It consists of a pipeline whose inputs are single cells and whose output is a lineage tree of input cells.

Humans

Paired Single-Cell Transcriptome and DNA Barcode Detection in Zebrafish Using ScarTrace.

ScarTrace is a CRISPR/Cas9-based genetic lineage tracing method that allows for uniquely barcoding the DNA of single cells at a target GFP sequence during developing zebrafish embryos. Single cells from barcoded adult zebrafish can be isolated from various tissues (e.g., marrow, brain, eyes, fins), and their transcriptome and barcode sequences are captured by single-cell cDNA amplification and genomic DNA nested PCR, respectively. Computationally, cell type and barcode identification permit clone tracing and lineage tree reconstruction of tissues to unravel fate decisions during embryogenesis.

Animals

Cloning and validating systems for high throughput molecular recording.

Molecular recording technologies record and store information about cellular history. Lineage tracing is one form of molecular recording and produces information describing cellular trajectories during mammalian development, differentiation and maintenance of adult stem cell niches, and tumor evolution. Our molecular recorder technology utilizes CRISPR-Cas9 barcode editing to generate mutations in genomically integrated, engineered DNA cassettes, which are read out by single-cell RNA sequencing and used to produce high-resolution lineage trees. Here, we describe optimized cloning and validation procedures to construct the molecular recorder lineage tracing system. We include information on considerations of technology design, cloning procedures, the generation of lineage tracing cell lines, and time course experiments to assess their performance.

Cloning, Molecular

Cell lineage, cell death, and the developmental origin of identified serotonin- and dopamine-containing neurons in the leech.

The nervous system of the glossiphoniid leech includes segmentally iterated neurons that contain serotonin (5-HT) and dopamine. These have been investigated in Helobdella triserialis, Theromyzon rude, and Haementeria ghilianii. Five types of 5-HT neurons are identified by immunocytochemistry in the abdominal ganglia of the ventral nerve cord: the bilaterally paired Retzius, anteromedial, ventrolateral and dorsolateral neurons, and the unpaired posteromedial (pm) neuron. Three types of bilaterally paired dopamine neurons are identified by glyoxylic acid-induced fluorescence in the segmental body wall: MD, LD1, and LD2. Each left or right half of the segmental complement of the leech nervous system is known to develop from 6 distinct ectodermal primary blast cells (ns, nf, o, p, qs, and qf). To identify the blast cells of origin of the 5-HT and dopamine neurons, fluorescent cell lineage tracers were injected into the various precursors of the blast cells in early (stage 6) embryos. The embryos were then raised until their 5-HT and dopamine neurons could be scored (stage 11) for the presence or absence of lineage tracer. We find that the Retzius, anteromedial, and posteromedial 5-HT neurons are derived from the ns blast cell, while the ventrolateral and dorsolateral 5-HT neurons are derived from the nf blast cell. The unpaired pm 5-HT neuron arises as one of a bilateral pair of neurons, of which one later dies. Whether the left or right pm neuron survives in any given ganglion is the consequence of some form of competitive interaction between cells derived from the left and right n primary blast cells, possibly between the left and right pm neurons themselves. We find that, of the dopamine neurons, the LD1 neuron is derived from the o blast cell, the LD2 neurons from the p blast cell, and the MD neuron from one of the 2 kinds of q blast cells. These results show that the 5-HT and dopamine neurons arise from 5 different primary blast cells in a highly determinate manner, and they support the view that cells of a similar phenotype need not be closely related in the developmental cell lineage tree.

Animals

Phylogenetic scanning: a computer-assisted algorithm for mapping gene conversions and other recombinational events.

An algorithm, 'phylogenetic scanning', is described for mapping gene conversion events where comparative DNA sequence data are available from different species. In this algorithm, sets of hypothetical phylogenetic trees are constructed that describe possible sequence relationships due to gene conversions in different species lineages; these trees are then evaluated by the principle of parsimony at intervals in the sequence alignment. When used to map gene conversion events that occurred between the pair of gamma-globin genes of higher primates, the algorithm gives results nearly identical to those obtained using a tedious manual approach. Suggestions are also provided for adaptation of this procedure to the analysis of other recombination events.

Algorithms

A framework for automated scalable designation of viral pathogen lineages from genomic data.

Pathogen lineage nomenclature systems are a key component of effective communication and collaboration for researchers and public health workers. Since February 2021, the Pango dynamic lineage nomenclature for SARS-CoV-2 has been sustained by crowdsourced lineage proposals as new isolates were sequenced. This approach is vulnerable to time-critical delays as well as regional and personal bias. Here we developed a simple heuristic approach for dividing phylogenetic trees into lineages, including the prioritization of key mutations or genes. Our implementation is efficient on extremely large phylogenetic trees consisting of millions of sequences and produces similar results to existing manually curated lineage designations when applied to SARS-CoV-2 and other viruses including chikungunya virus, Venezuelan equine encephalitis virus complex and Zika virus. This method offers a simple, automated and consistent approach to pathogen nomenclature that can assist researchers in developing and maintaining phylogeny-based classifications in the face of ever-increasing genomic datasets.

Animals

On the origin of animals and placental mammals: a critique of literalist readings of the fossil record.

The fossil record is incomplete, as evidenced by the pervasive presence of ghost lineages throughout the Tree of Life. For example, across placental mammals, at least 720 Myr of basal lineages are ghost lineages, that is, lineages that have left no fossil evidence of their past history. In contrast, some studies have suggested that the fossil record is a faithful temporal archive of evolutionary history and thus the times of diversification of clades must be close to the ages of their oldest fossils. Such literalist interpretations have been contradicted by analysis of molecular datasets which, in many cases, indicate that groups including placental mammals and animals may have originated at times substantially older than their fossil records. Some of those studies have further argued that, in the case of animals and placental mammals, molecular clocks are uninformative, suffer from characteristic pathologies, and thus cannot distinguish between recent and ancient hypotheses of diversification. Here, we reexamine these two cases and show, using Bayesian model selection theory, that the explosive diversification models previously proposed for animals and placental mammals have a posterior probability of ∼0. We show the characteristic pathologies purportedly discovered do not exist, highlight errors in previous analyses, and provide advice on best practice for molecular-clock dating analysis.

Animals

Phylogenetic analysis based on rRNA sequences supports the archaebacterial rather than the eocyte tree.

How many primary lineages of life exist and what are their evolutionary relationships? These are fundamental but highly controversial issues. Woese and co-workers propose that archaebacteria, eubacteria and eukaryotes are the three primary lines of descent and their relationships can be represented by Fig. 1a (the 'archaebacterial tree') if one neglects the root of the tree. In contrast, Lake claims that archaebacteria are paraphyletic, and he groups eocytes (extremely thermophilic, sulphur-dependent bacteria) with eukaryotes, and halobacteria with eubacteria (the 'eocyte tree', Fig. 1b). Lake's view has gained considerable support as a result of an analysis of small subunit ribosomal RNA sequence data by a new approach, the evolutionary parsimony method. Here we report that analysis of small subunit data by the neighbour-joining and maximum parasimony methods favours the archaebacterial tree and that computer simulations using either the archaebacterial or the eocyte tree as a model tree show that the probability of recovering the model tree is very high (greater than 90 per cent) for both the neighbour-joining and maximum parsimony methods but is relatively low for the evolutionary parsimony method. Moreover, analysis of large subunit rRNA sequences by all three methods strongly favours the archaebacterial tree.

Archaea

Statistical method for estimating the standard errors of branch lengths in a phylogenetic tree reconstructed without assuming equal rates of nucleotide substitution among different lineages.

A statistical method is developed for estimating the standard errors of branch lengths in a phylogenetic tree reconstructed without assuming equal rates of nucleotide substitution among different lineages. This method can be easily used for testing whether the length of an interior branch in a reconstructed tree is positive, i.e., whether the topology of the tree is correct. Computer simulations indicate that this method is appropriate for a statistical test. As an example, this method is applied to phylogenetic trees reconstructed for the four hominoid species: human, chimpanzee, gorilla, and orangutan. The results obtained show that the present method provides a powerful statistical test.

Animals

Unraveling evolutionary pathways: allopolyploidization and introgression in polyploid Prunus (Rosaceae).

Allopolyploidization, resulting from hybridization and subsequent whole-genome duplication (WGD), is a fundamental mechanism driving evolutionary diversification across various lineages within the Tree of Life. The polyploid Prunus (Rosaceae), significant for its economic and agricultural value, provides an ideal model for investigating the evolutionary dynamics associated with allopolyploidy. In this study, we utilized deep genome skimming (DGS) data to demonstrate a comprehensive analytical framework for elucidating the underlying allopolyploidy that includes a newly adapted tool (DGS-Tree2GD) tailored explicitly for accurately detecting WGD events. Additionally, we introduced two methods to evaluate the contribution of incomplete lineage sorting (ILS) to lineage diversification. Phylogenomic discordance analyses revealed that allopolyploidization, rather than ILS, played a dominant role in the origin and dynamics of polyploid Prunus. Moreover, we inferred that the uplift of the Himalayas from the Middle to Late Miocene was a key driver in the rapid diversification of the Maddenia clade, an endemic group in East Asia. This geological event facilitated extensive hybridization and allopolyploidization, particularly the introgression between the Himalayas-Hengduan and Central-Eastern China clades. This case study demonstrates the robustness and efficacy of our analytical approach in precisely identifying WGD events and elucidating the evolutionary mechanisms underlying allopolyploidization in polyploid Prunus.

Polyploidy

Gene phylogenies and the endosymbiotic origin of plastids.

The endosymbiotic origin of chloroplasts from cyanobacteria has long been suspected and has been confirmed in recent years by many lines of evidence. Debate now is centered on whether plastids are derived from a single endosymbiotic event or from multiple events involving several photosynthetic prokaryotes and/or eukaryotes. Phylogenetic analysis was undertaken using the inferred amino acid sequences from the genes psbA, rbcL, rbcS, tufA and atpB and a published analysis (Douglas and Turner, 1991) of nucleotide sequences of small subunit (SSU) rRNA to examine the relationships among purple bacteria, cyanobacteria and the plastids of non-green algae (including rhodophytes, chromophytes, a cryptophyte and a glaucophyte), green algae, euglenoids and land plants. Relationships within and among groups are generally consistent among all the trees; for example, prochlorophytes cluster with cyanobacteria (and not with green plastids) in each of the trees and rhodophytes are ancestral to or the sister group of the chromophyte algae. One notable exception is that Euglenophytes are associated with the green plastid lineage in psbA, rbcL, rbcS and tufA trees and with the non-green plastid lineage in SSU rRNA trees. Analysis of psbA, tufA, atpB and SSU rRNA sequences suggests that only a single bacterial endosympbiotic event occurred leading to plastids in the various algal and plant lineages. In contrast, analysis of rbcL and rbcS sequences strongly suggests that plastids are polyphyletic in origin, with plastids being derived independently from both purple bacteria and cyanobacteria. A hypothesis consistent with these discordant trees is that a single bacterial endosymbiotic event occurred leading to all plastids, followed by the lateral transfer of the rbcLS operon from a purple bacterium to a rhodophyte.

Amino Acid Sequence

Phylotranscriptomics Allows Distinguishing Major Gene Flow Events from Incomplete Lineage Sorting in Rapidly Diversifying Mimetic Orchids (Genus Ophrys).

Ophrys orchids (or bee orchids) provide an outstanding example of a plant adaptive radiation. Over the last 5 million years, this genus has diversified into hundreds of taxa as a result of its unconventional pollination strategy, known as "sexual swindling". However, the rapid and substantial diversification of this genus, combined with its capacity for hybridization and large genome size, poses significant challenges in addressing its systematics. We used phylotranscriptomics as a genome complexity reduction technique to infer the phylogenetic relationships among Ophrys main lineages. More than seven thousand gene trees enabled us to determine the relative contributions of gene flow and incomplete lineage sorting (ILS) in Ophrys evolution. First, we propose a new phylogenetic hypothesis for the genus with an unprecedented resolution that largely confirms the relationships between the main Ophrys lineages, but also provides new insights within each subgenera. By combining phylogenetic network inference with introgression analyzes based on gene tree topologies and branch lengths, we then show that the numerous phylogenetic incongruences among gene tree topologies result from a pervasive background of ILS, over which stand out several well-supported, ancient and potentially adaptive gene flow events between lineages. These major gene flow events provide a new perspective on the evolution of the Ophrys genus and its pollination, questioning previous hypotheses inferred without considering its reticulate evolution, and providing a better understanding of discrepancies observed among previous phylogenetic studies of the genus.

Orchidaceae

Evolution of the regulatory isozymes of 3-deoxy-D-arabino-heptulosonate 7-phosphate synthase present in the Escherichia coli genealogy.

The evolutionary history of isozymes for 3-deoxy-D-arabino-heptulosonate 7-phosphate (DAHP) synthase has been constructed in a phylogenetic cluster of procaryotes (superfamily B) that includes Escherichia coli. Members of superfamily B that have been positioned on a phylogenetic tree by oligonucleotide cataloging possess one or more of four distinct isozymes of DAHP synthase. DAHP synthase-0 is insensitive to feedback inhibition, while DAHP synthase-Tyr, DAHP synthase-Trp, and DAHP synthase-Phe are sensitive to feedback inhibition by L-tyrosine, L-tryptophan, and L-phenylalanine, respectively. The evolutionary history of this isozyme family can be deduced within superfamily B by using a cladistic methodology of maximum parsimony (R. A. Jensen, Mol. Biol. Evol. 2:92-108, 1985). DAHP synthase-0 was found in Acinetobacter species and in Oceanospirillum minutulum, organisms that also possess DAHP synthase-Tyr. These two isozymes were apparently present in a common ancestor that predated the evolutionary divergence of contemporary superfamily B sublineages. DAHP synthase-0 is postulated to have been the evolutionary forerunner of DAHP synthase-Trp. The newly evolved DAHP synthase-Trp is postulated to have possessed sensitivity to feedback inhibition by chorismate as well as by L-tryptophan, chorismate sensitivity having been retained in rRNA group I pseudomonads (minor sensitivity), group V pseudomonads (very sensitive), and Lysobacter enzymogenes (ultrasensitive). Organisms constituting the enteric lineage of the phylogenetic tree (including a cluster of four Oceanospirillum species) have all lost the chorismate sensitivity of DAHP synthase-Trp. The absence of DAHP synthase-Phe in the Oceanospirillum cluster of organisms supports the previous conclusion that DAHP synthase-Phe evolved recently within superfamily B, being present only Escherichia coli and its close relatives.

3-Deoxy-7-Phosphoheptulonate Synthase

Evolution of the sarafotoxin/endothelin superfamily of proteins.

Sixteen protein and nucleic acid sequences from the vasoconstrictor sarafotoxin/endothelin/endothelin-like superfamily of peptides were studied, and the evolutionary relationships between the sarafotoxin and endothelin gene families as well as the phylogenetic topology within each gene family and the three endothelin subfamilies was reconstructed. The endothelin gene family has diverged from an ancestral gene that has experienced an exon duplication event followed by two gene duplication events. The sarafotoxins' lineage diverged from the ancestral gene prior to the first endothelin gene duplication event. Analysis of the resulting phylogenetic trees revealed that in several lineages, the peptides have independently accumulated identical replacements in position 2, therefore supporting the hypothesis that residue 2 is crucial to their activity.

Amino Acid Sequence