Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Efficient extension of a misaligned tRNA-primer during replication of the HIV-1 retrovirus.

The human immunodeficiency virus (HIV) and other retroviruses show extensive genomic variation, which is primarily due to error-prone replication by the viral reverse transcriptase (RT) enzymes. RT errors include misincorporation with subsequent extension of the mismatched terminal base, and extension of realigned primer-template duplexes. Whereas both RT-mediated mechanisms have been extensively studied in vitro, almost no in vivo experiments have been performed. In this work, we analyzed the ability of HIV-1 RT to extend a misaligned tRNA(Lys3) primer in vivo. This tRNA binds with its 3'-terminal 18 nt to a complementary sequence in the viral genome, referred to as the primer-binding site (PBS). We constructed a series of mutant viral genomes with small insertions or deletions in the PBS sequence, resulting in misalignment of the tRNA primer. Extension of the misaligned primer did occur with reasonable efficiency for some of the mutants, resulting in reversion to the wild-type viral sequence. The infectivity and reversion frequency of the PBS mutants is therefore a measure of the efficiency of extending a misaligned primer in vivo. Using virion-derived primer-template complexes, we also measured the tRNA-priming efficiency in vitro. The combined results show that HIV-1 RT can elongate a misaligned primer and that the efficiency of primer extension is determined by the extent of the mismatch.

Base Sequence↗

Quasispecies in retrotransposons: a role for sequence variability in Tnt1 evolution.

Retroviral replication is a very error-prone process. Replication of retroviruses gives rise to populations of closely related but different genomes referred to as 'quasispecies'. This huge swarm of different sequences constitutes a reservoir of potentially useful genomes in case of an environmental change, endowing retroviruses with extreme adaptability. Retrotransposons are mobile genetic elements closely related to retroviruses, and retrotransposition is as error prone as retroviral replication. The Tnt1 retrotransposon is present in hundreds of copies in the genome of tobacco that show a high level of sequence heterogeneity. When Tnt1 is expressed, its RNA is not a single sequence but a population of sequences displaying a quasispecies-like structure. This population structure gives to Tnt1, as in the case of retroviruses, a high sequence plasticity and an adaptive capacity. We propose this adaptivity as the major reason for Tnt1 maintenance in Nicotiana genomes and we discuss in this paper the importance of sequence variability for Tnt1 evolution.

Adaptation, Physiological↗

Encapsidation of poliovirus replicons encoding the complete human immunodeficiency virus type 1 gag gene by using a complementation system which provides the P1 capsid protein in trans.

Poliovirus genomes which contain small regions of the human immunodeficiency virus type 1 (HIV-1) gag, pol, and env genes substituted in frame for the P1 capsid region replicate and express HIV-1 proteins as fusion proteins with the P1 capsid precursor protein upon transfection into cells (W. S. Choi, R. Pal-Ghosh, and C. D. Morrow, J. Virol. 65:2875-2883, 1991). Since these genomes, referred to as replicons, do not express capsid proteins, a complementation system was developed to encapsidate the genomes by providing P1 capsid proteins in trans from a recombinant vaccinia virus, VV-P1. Virus stocks of encapsidated replicons were generated after serial passage of the replicon genomes into cells previously infected with VV-P1 (D. C. Porter, D. C. Ansardi, W. S. Choi, and C. D. Morrow, J. Virol. 67:3712-3719, 1993). Using this system, we have further defined the role of the P1 region in viral protein expression and RNA encapsidation. In the present study, we constructed poliovirus replicons which contain the complete 1,492-bp gag gene of HIV-1 substituted for the entire P1 region of poliovirus. To investigate whether the VP4 coding region was required for the replication and encapsidation of poliovirus RNA, a second replicon in which the complete gag gene was substituted for the VP2, VP3, and VP1 capsid sequences was constructed. Transfection of replicon RNA with and without the VP4 coding region into cells resulted in similar levels of expression of the HIV-1 Gag protein and poliovirus 3CD protein, as indicated by immunoprecipitation using specific antibodies. Northern (RNA) blot analysis of RNA from transfected cells demonstrated comparable levels of RNA replication for each replicon. Transfection of the replicon genomes into cells infected with VV-P1 resulted in the encapsidation of the genomes; serial passage in the presence of VV-P1 resulted in the generation of virus stocks of encapsidated replicons. Analysis of the levels of protein expression and encapsidated replicon RNA from virus stocks after 21 serial passages of the replicon genomes with VV-P1 indicated that the replicon which contained the VP4 coding region was present at a higher level than the replicon which contained a complete substitution of the P1 capsid sequences. These differences in encapsidation, though, were not detected after only two serial passages of the replicons with VV-P1 or upon coinfection and serial passage with type 1 Sabin poliovirus.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Chromosome-level assembly and annotation of the Jaguar (Panthera onca) genome.

OBJECTIVES: The Jaguar (Panthera onca) is a large cat species native to the Americas. Despite being successful predators, jaguar populations have declined due to habitat loss. Genome resources can help in conservation efforts as well as in understanding the interesting biology of these Felids. Beside contiguity, a well annotated reference genome provides contextual information for variants that will benefit the design of appropriate conservation programs. DATA DESCRIPTION: We sequenced material from two individuals using a combination of ONT reads and Illumina PE. The resulting nuclear genome assembly has a larger contig N50 (48.04 Mb) compared with the existing annotated chromosome-level assembly published by the DNA Zoo project. Using public Hi-C data, we obtained an improved chromosome-level assembly of the Jaguar genome (mPanOnc3.5) with larger contigs, 99.85% of the sequence assigned to chromosomes and 25,267 protein coding genes annotated. Overall, this improved assembly provides a better reference to study this threatened species.

Animals↗

A phylogenetic analysis by multilocus enzyme electrophoresis and multiprimer random amplified polymorphic DNA fingerprinting of the Leishmania genome project Friedlin reference strain.

We have assessed the phylogenetic status of the Leishmania genome project Friedlin reference strain by MLEE and multiprimer RAPD including a set of 9 stocks representative of the main Leishmania species and of the whole genetic diversity of the Leishmania genus. To our knowledge, the detailed genetic characterization of the Friedlin strain has never been published before. As previously recorded (Tibayrenc et al. 1993), MLEE and RAPD data gave congruent phylogenetic results. The Friedlin reference strain was definitely attributed to Leishmania (Leishmania) major Yakimoff et Schokhor, 1914. Five specific RAPD patterns made it possible to distinguish between the Friedlin strain and the 2 other L. (L.) major stocks included in the study. Various specific MLEE and RAPD characters permitted to distinguish between the Leishmania species included in the study. All these characters are usable to detect accidental laboratory mix-ups involving the Friedlin reference strain. In confirmation with previous studies involving a more limited set of genetic markers, the general genetic diversity of the Leishmania genus proved to be considerable. It must be made clear that only one strain cannot be considered as representative of the whole genetic variability of the genus Leishmania. In the future, it is therefore advisable to complement the results obtained in the framework of the Leishmania genome project with data from other strains that should be selected on a criterion of important genetic differences with the Friedlin strain.

Animals↗

Quantitative studies of integration of murine leukemia virus after exogenous infection.

Using a [3H]DNA probe prepared from AKR murine leukemia virus, we determined the number of copies of the AKR virus genome integrated into the cellular DNA after exogenous infection of NIH mouse, AKR mouse, and rat cells in tissue culture. NIH mouse cells, which lack a portion of the viral genome (referred to as Gross-AKR specific sequences), incorporated three to four copies of these sequences per haploid genome. AKR cells, in which the Gross-AKR specific sequences are already present as three to four copies per haploid genome, did not shwo any distinct change in copy number after infection. Rat cells, which lack DNA sequences homologous to murine leukemia virus, incorporated one copy of the viral genome per haploid genome. It is inferred that the presence of viral sequences may affect the efficiency of integration of exogenous provirus, and that there may be a limit to the number of copies that can be inserted.

AKR murine leukemia virus↗

High-throughput analysis of subtelomeric chromosome rearrangements by use of array-based comparative genomic hybridization.

Telomeric chromosome rearrangements may cause mental retardation, congenital anomalies, and miscarriages. Automated detection of subtle deletions or duplications involving telomeres is essential for high-throughput diagnosis, but impossible when conventional cytogenetic methods are used. Array-based comparative genomic hybridization (CGH) allows high-resolution screening of copy number abnormalities by hybridizing differentially labeled test and reference genomes to arrays of robotically spotted clones. To assess the applicability of this technique in the diagnosis of (sub)telomeric imbalances, we here describe a blinded study, in which DNA from 20 patients with known cytogenetic abnormalities involving one or more telomeres was hybridized to an array containing a validated set of human-chromosome-specific (sub)telomere probes. Single-copy-number gains and losses were accurately detected on these arrays, and an excellent concordance between the original cytogenetic diagnosis and the array-based CGH diagnosis was obtained by use of a single hybridization. In addition to the previously identified cytogenetic changes, array-based CGH revealed additional telomere rearrangements in 3 of the 20 patients studied. The robustness and simplicity of this array-based telomere copy-number screening make it highly suited for introduction into the clinic as a rapid and sensitive automated diagnostic procedure.

Chromosome Aberrations↗

Multi-species comparative mapping in silico using the COMPASS strategy.

MOTIVATION: The completion of human and mouse genome sequences provides a valuable resource for decoding other mammalian genomes. The comparative mapping by annotation and sequence similarity (COMPASS) strategy takes advantage of the resource and has been used in several genome-mapping projects. It uses existing comparative genome maps based on conserved regions to predict map locations of a sequence. An automated multiple-species COMPASS tool can facilitate in the genome sequencing effort and comparative genomics study of other mammalian species. RESULTS: The prerequisite of COMPASS is a comparative map table between the reference genome and the predicting genome. We have built and collected comparative maps among five species including human, cattle, pig, mouse and rat. Cattle-human and pig-human comparative maps were built based on the positions of orthologous markers and the conserved synteny groups between human and cattle and human and pig genomes, respectively. Mouse-human and rat-human comparative maps were based on the conserved sequence segments between the two genomes. With a match to human genome sequences, the approximate location of a query sequence can be predicted in cattle, pig, mouse and rat genomes based on the position of the match relatively to the orthologous markers or the conserved segments. AVAILABILITY: The COMPASS-tool and databases are available at http://titan.biotec.uiuc.edu/COMPASS/

Algorithms↗

Identification and characterization of human FMNL1, FMNL2 and FMNL3 genes in silico.

FMNL (NM_005892.2) is a 5'-truncated partial cDNA encoding a Formin-homology protein related to DAAM1, DAAM2, DIAPH1 and DIAPH2. Here, we identified three members of FMNL gene family in the human genome by using bioinformatics. FMNL1 gene, corresponding to 5'-truncated KW-13 and FMNL cDNAs, was located within reference genomic contig NT_010748.9 (nucleotide position 100576-125849, forward orientation). FMNL2 gene, corresponding to KIAA1902 and FHOD2 cDNAs, was located within NT_005151.10 (nucleotide position 122465-436828, forward orientation). FMNL3 gene, corresponding to 5'-truncated DKFZp762B245 and KIAA2014 cDNAs, was located within NT_026397.10 (nucleotide position 209769-279037, reverse orientation). FMNL1, FMNL2 and FMNL3 genes encode A and B isoforms with the C-terminal divergence due to alternative splicing (cassette splicing of exon 26). FMNL1A (1100 aa), FMNL1B (1114 aa), FMNL2A (1087 aa), FMNL2B (1093 aa), FMNL3A (1028 aa) and FMNL3B (1027 aa) consist of FDD, FH1 and FH2 domains. Total amino-acid identity were as follows: FMNL1A vs. FMNL2A, 59.3%; FMNL1A vs. FMNL3A, 56.1%; FMNL2A vs. FMNL3A, 68.6%. FMNL1 gene was mapped to human chromosome 17q21. FMNL2 gene was linked to FNBP3/HYPA gene on chromosome 2q23.3, while FMNL3 gene was linked to FNBP3L/HYPC gene on chromosome 12q13. FMNL1 mRNA was expressed in natural killer cells, Burkitt lymphoma, pancreatic cancer, prostate cancer, and lung large cell carcinoma, FMNL2 mRNA in several normal tissues, diffuse-type gastric cancer, breast cancer, chondrosarcoma, melanoma, and glioblastoma, and FMNL3 mRNA in gastric cancer. FMNL1, FMNL2 and FMNL3 might be implicated in polarity control, invasion, migration, or metastasis through regulation of the Rho-related signaling pathway.

Alternative Splicing↗

T2T genomes of Caenorhabditis nigoni and Caenorhabditis briggsae reveal divergence in satellite DNA abundance.

The two closely related nematode species, Caenorhabditis nigoni and Caenorhabditis briggsae, are commonly used to study the evolution of reproductive modes in animals, with the self-fertile C. briggsae and outcrossing C. nigoni sharing a common ancestor ∼3.5 million years ago. Earlier genomic analyses revealed that selfing Caenorhabditis species have smaller genomes and proposed that at least some gene loss in C. briggsae is adaptive. However, the incomplete C. nigoni reference genome has limited most comparative analyses to genic regions. Here, we leverage long-read sequencing to generate and annotate telomere-to-telomere (T2T) assemblies for the C. nigoni strain JU1422 and the C. briggsae strain AF16. This new 139 Mb C. nigoni genome resolves 57 gaps and 149 unassigned scaffolds from the previous genome assembly. A major driver of the size difference with the 107 Mb T2T C. briggsae genome is the abundance of satellite DNA, which accounts for 12.8 Mb (9.2%) in C. nigoni and only 3.2 Mb (3.0%) in C. briggsae Notably, the C. nigoni X Chromosome is 13.4 Mb larger than in the previous assembly, making it 60% larger than the C. briggsae X Chromosome compared with 18%-26% difference for the autosomes. We also document a surprising degree of plasticity in the ribosomal DNA, with the C. nigoni X Chromosome harboring a second 45S rDNA array that is absent in C. briggsae The hitherto undocumented divergence in the abundance of repetitive DNA elements makes the new genomes an invaluable resource for genomic analysis.

Journal Article↗

Whole genome sequence data on Ethiopian key sorghum landraces and founder lines.

Sorghum (Sorghum bicolor (L.) Moench) is the fifth most important cereal globally. Its genetic diversity is key to improving yield stability, stress tolerance, and adaptation to different environments. Ethiopia is one of the centers for the crop's origin, diversity, and use in both human food and livestock feed. However, genomic data on Ethiopian sorghum remain limited, especially for landraces preferred by local farmers. This dataset consists of whole-genome sequencing data for 188 Ethiopian sorghum accessions, including founder lines and important landraces from major agroecological zones. Sequencing was performed using the Complete Genomics DNBSEQ-T7 platform, generating high-coverage whole-genome data (20 &#xd7; coverage). On average, each accession produced 64.7 million reads. Reads were aligned to the Sorghum bicolor NCBIv3 reference genome and variants called using GATK HaplotypeCaller with joint genotyping (GATK v4.6.1.0). Hard-filtering followed GATK best-practice thresholds (QD <2.0, FS> 60.0, MQ <40.0, MQRankSum <-12.5, ReadPosRankSum <-8.0), retaining biallelic SNPs with mean depth 10-50&#xd7;, missingness &#x2264;20%, and MAF &#x2265;0.05, yielding 6095,752 high-confidence SNPs across 185 accessions. Both the raw FASTQ files and processed VCF files are publicly available to support studies of sorghum genetic diversity, population structure, selection, and the genetic basis of important traits.

Adaptation↗

A Chromosome-Level Genome Assembly of the Potato Leafhopper Empoasca fabae (Hemiptera: Cicadellidae).

The potato leafhopper, Empoasca fabae (Harris, 1841), is a highly polyphagous, migratory insect pest of eastern North America that feeds on more than 200 herbaceous and woody plant species, causing substantial losses to forage and field crops. Despite its agricultural and ecological importance, no genome has been available for this species. Here, we present the first chromosome-level genome assembly of E. fabae, generated from Oxford Nanopore long reads, Illumina short reads, and Omni-C proximity-ligation data. The final assembly spans 908&#x2005;Mb across 132 scaffolds, with 99.8% of the assembly captured in ten chromosome-length scaffolds (nine autosomes and an X chromosome) with a scaffold N50 of 96.2&#x2005;Mb. The assembly is highly complete, recovering 92.9% of conserved hemipteran single-copy orthologs from protein annotations, and is composed of 47.6% repetitive sequence, dominated by long terminal repeat retrotransposons and unclassified elements. Read-depth comparison between male and female individuals supports assignment of a single sex-linked chromosome, consistent with an XO sex determination system. BRAKER3 gene annotation predicted 31,406 protein-coding genes after retaining the longest isoform per locus. Comparative genome analysis of the two closest related Typhlocybinae species with genomes available, Matsumurasca onukii and Hebata decipiens, revealed extensive chromosome-scale collinearity while defining a shared core gene repertoire. This reference genome provides a foundation for comparative and population genomic studies and for investigating genetic traits in this economically important crop pest species.

Animals↗

Segmental duplications: organization and impact within the current human genome project assembly.

Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. Initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and > or =1 kb in length). Duplications (90%-98%; > or =1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications >98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.

Base Sequence↗

T2T genomes of Caenorhabditis nigoni and Caenorhabditis briggsae reveals extensive loss of satellite DNA associated with self-fertilization.

The two closely related Caenorhabditis nematode species, C. nigoni and C. briggsae , are commonly used to study the evolution of reproductive modes in animals, with the self-fertile C. briggsae and outcrossing C. nigoni sharing a common ancestor &#x223c;3.5 million years ago. Earlier genomic analyses of these species revealed genome shrinkage associated with selfing and proposed that at least some gene loss can be adaptive. However, the incomplete C. nigoni reference genome limited most comparative analyses to genic regions. Here, we leveraged long-read sequencing to generate a telomere-to-telomere (T2T) assembly for the C. nigoni strain JU1422 and the C. briggsae strain AF16. This new 139Mb C. nigoni genome resolved 57 gaps and 149 unassigned scaffolds from the previous genome assembly. Comparison with the 107Mb T2T C. briggsae genome reveals that the major driver of genome content differences are deletions to satellite DNA arrays, reflecting a loss of 9.6Mb. Interestingly, many of the differences are on the C. nigoni X chromosome, which is >13Mb larger than in the previous assembly. The transition to selfing was thus accompanied by a 37% reduction in the size of the sex chromosome compared to 16-21% shrinkage of the autosomes. We also document a surprising degree of plasticity in the ribosomal DNA, with the X chromosome harboring a second 45S rDNA array that is absent in C. briggsae . Our analysis reveals that obligatory outcrossing may play a major role in the maintenance of satellite DNA arrays.

Journal Article↗

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256&#x2009;Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms↗

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55&#x2009;Gb, with a scaffold N50 of 93.38&#x2009;Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant↗

A chromosome-scale genome of Capsicum pubescens provides insights into candidate terpene-associated gene clusters and pan variation of terpene synthases.

A chromosome-scale genome of Capsicum pubescens and comparative pan-TPS analysis support structural characterization and gene-level prioritization of a chromosome-9 terpene-associated candidate locus in this accession. Capsicum pubescens is one of the five domesticated Capsicum species, mainly cultivated in mid- to high-elevation regions of the Americas. Despite its distinctive morphology and fruit traits, genomic resources for C. pubescens remain less developed than those for the widely cultivated C. annuum. Here, we assembled a chromosome-scale reference genome for accession HNUCP0001, spanning 3.70&#xa0;Gb with a scaffold N50 of 278.01&#xa0;Mb. Comparative genomics revealed 679 significantly expanded gene families enriched in sesquiterpenoid and triterpenoid biosynthesis. Genome-wide biosynthetic gene-cluster mining identified multiple terpene-associated candidate loci, which were subsequently prioritized using genome-derived structural criteria and Capsicum pubescens-specific expression evidence. Subsequently, we curated the terpene synthase (TPS) repertoire and, across 16 Capsicum genomes, resolved 36 TPS orthogroups with pronounced presence/absence variation, highlighting dynamic lineage-specific diversification. Together, these analyses establish HNUCP0001 as an accession-specific genomic resource and provide a comparative framework for prioritizing terpene-associated TPS genes and candidate BGCs in Capsicum. These candidate loci, together with accession-level transcriptomic and metabolomic evidence, offer testable hypotheses for future functional studies of specialized terpenoid metabolism in C. pubescens.

Alkyl and Aryl Transferases↗

The genome sequence of Sphagnum contortum Schultz, 1819 (Sphagnales: Sphagnaceae).

We present a genome assembly of Sphagnum contortum (twisted bog-moss; Streptophyta; Sphagnopsida; Sphagnales; Sphagnaceae). The genome sequence has a total length of 370.55 megabases. Most of the assembly (99.46%) is scaffolded into 21 chromosomal pseudomolecules. The mitochondrial sequence has a length of 141.7 kilobases and the plastid genome assembly has a length of 140.11 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Sphagnales↗