Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

The genome sequence of Sphagnum tenellum (Brid.) Bory, 1819 (Sphagnales: Sphagnaceae).

We present a genome assembly of Sphagnum tenellum (soft bog-moss; Streptophyta; Sphagnopsida; Sphagnales; Sphagnaceae). The genome sequence has a total length of 386.16 megabases. Most of the assembly (98.04%) is scaffolded into 21 chromosomal pseudomolecules. The mitochondrial sequence has a length of 141.31 kilobases and the plastid genome assembly has a length of 140.16 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Sphagnales↗

The genome sequence of Carex sylvatica Huds., 1762 (Poales: Cyperaceae).

We present a genome assembly of Carex sylvatica (Wood-sedge; Streptophyta; Magnoliopsida; Poales; Cyperaceae). The genome sequence has a total length of 363.24 megabases. Most of the assembly (99.79%) is scaffolded into 29 chromosomal pseudomolecules. The mitochondrial sequences have lengths of 1715.64, 74.78 and 77.37 kilobases and the plastid genome assembly has a length of 215.14 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Carex sylvatica↗

The genome sequence of Prunus padus L., 1753 (Rosales: Rosaceae).

We present a genome assembly of Prunus padus (bird cherry; Streptophyta; Magnoliopsida; Rosales; Rosaceae). The genome sequence has a total length of 454.49 megabases. Most of the assembly (97.36%) is scaffolded into 16 chromosomal pseudomolecules. The mitochondrial sequences have lengths of 280.28 and 155.92 kilobases and the plastid genome assembly has a length of 158.96 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Prunus padus↗

The genome sequence of Hypericum androsaemum L., 1753 (Malpighiales: Hypericaceae).

We present a genome assembly of Hypericum androsaemum (Tutsan; Streptophyta; Magnoliopsida; Malpighiales; Hypericaceae). The genome sequence has a total length of 384.54 megabases. Most of the assembly (99.96%) is scaffolded into 20 chromosomal pseudomolecules. The mitochondrial sequences have lengths of 138.63, 75.57, 98.15 and 103.38 kilobases and the plastid genome assembly has a length of 165.0 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Hypericum androsaemum↗

The genome sequence of Hyocomium armoricum (Brid.) Wijk & Margad. (Hypnales: Hypnaceae).

We present a genome assembly of Hyocomium armoricum (Flagellate Feather-moss; Streptophyta; Bryopsida; Hypnales; Hypnaceae). The genome sequence has a total length of 333.60 megabases. Most of the assembly (99.59%) is scaffolded into 11 chromosomal pseudomolecules. The mitochondrial sequence has a length of 104.43 kilobases and the plastid genome assembly has a length of 123.84 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Flagellate Feather-moss↗

The genome sequence of Frangula alnus Mill., 1768 (Rosales: Rhamnaceae).

We present a genome assembly of Frangula alnus (alder buckthorn; Streptophyta; Magnoliopsida; Rosales; Rhamnaceae). The assembly consists of two haplotypes with total lengths of 279.06 megabases and 281.25 megabases. Most of haplotype 1 (99.91%) is scaffolded into 10 chromosomal pseudomolecules. Most of haplotype 2 (99.91%) is scaffolded into 10 chromosomal pseudomolecules. The mitochondrial sequence has a length of 492.03 kilobases and the plastid genome assembly has a length of 161.23 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain and Ireland.

Frangula alnus↗

Chromosome-level genome assembly of the ornamental plant Alcea rosea.

Alcea rosea, a member of the Malvaceae family, is celebrated for its rich floral palette and global horticultural significance. Here, we present a high-quality reference genome for A. rosea, achieving a genome assembly size of 1.01 Gbp, with a Contig N50 length of 36.61 Mbp. The genome sequence was successfully mapped to 21 chromosomes, and the scaffold N50 length reached 52.57 Mbp, with a scaffold genome completeness of 99.6%. A total of 565.84 Mbp (comprising 56% of the genome) of repetitive sequences were identified, with transposable elements being predominant, particularly long terminal repeat (LTR) elements, which accounted for 48.44% of the genome. 51,436 genes were annotated. Among these predicted genes, the average gene length and coding sequence (CDS) length were 2739.92 bp and 1242.54 bp, respectively.

Genome, Plant↗

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software↗

Whole-genome sequencing of adenovirus 41 directly from wastewater using nested overlapping PCR and MinION.

Human adenovirus F41 (HAdV-F41) is one of the leading causes of children's acute gastroenteritis and was recently linked to an outbreak of severe acute hepatitis of unknown etiology among children during 2021 to 2022. While most evidence is based on clinical data, wastewater-based epidemiology offers a community-level approach to monitoring circulating strains and enhancing outbreak preparedness. In this study, we developed an overlapping amplicon-based whole-genome sequencing approach to directly detect HAdV-F41 from archived wastewater samples, using nested PCR with 13 primer sets. Archived wastewater samples were collected between 2021 and 2022 from three treatment plants in Seattle, USA. The viral load ranged from 1.2 &#xd7; 103 to 8.4 &#xd7; 103 genome copies per liter. The Oxford Nanopore platform was used for whole-genome sequencing. Complete or partial (>84%) HAdV-F41 genomes were recovered from wastewater samples, with mean coverage depths ranging from 10&#xb3; to 10&#x2075;. The consensus sequences showed more than 99% similarity to reference genomes in the NCBI database. The phylogenetic analysis revealed that 2 sequences clustered within lineage 2a and 11 within lineage 2b, reflecting that at least two sub-lineages were circulating in the community at that time. Our results demonstrate that the overlapping amplicon-based whole-genome sequencing approach using the Oxford Nanopore platform reliably recovers HAdV-F41 genomes from wastewater. This method offers high-resolution genomic surveillance of circulating, clinically relevant HAdV-F41, supporting wastewater-based epidemiology as a valuable tool for detecting emerging variants and strengthening the early warning system for future disease outbreaks.IMPORTANCEHuman adenovirus F41 is a primary cause of childhood gastroenteritis and has been linked to recent outbreaks of severe acute hepatitis in children, yet community-level genomic surveillance of this virus remains limited. This study shows that wastewater can be used to recover nearly complete HAdV-F41 genomes through a targeted overlapping-amplicon sequencing strategy on the Oxford Nanopore platform. By applying this method to archived wastewater samples, we detected the simultaneous circulation of multiple viral lineages in a large city. These findings extend wastewater-based epidemiology beyond SARS-CoV-2 and emphasize its importance for monitoring clinically significant enteric viruses. The method described here offers a scalable tool for tracking viral evolution in communities and enhancing early warning systems for future outbreaks.

Wastewater↗

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes↗

Characterisation of Trichuris incognita n sp in C&#xf4;te d'Ivoire: a morphological, genomic, and genome-wide association with drug sensitivity study.

BACKGROUND: Trichuriasis is a neglected tropical disease that affects up to 500 million individuals and can cause considerable morbidity. For decades, trichuriasis was thought to be caused by one species of whipworm, Trichuris trichiura. The aim of this study was to investigate the origin of differences in response rates to the best available anthelmintic treatment for trichuriasis-a combination of albendazole and ivermectin-in C&#xf4;te d'Ivoire by analysing the parasite population. METHODS: In this morphological, genomic, and genome-wide association study (GWAS) with drug sensitivity we used long-read and short-read sequencing approaches and assembled a high-quality reference genome of Trichuris incognita n sp isolated in a primary interventional study conducted in the Lagunes district of C&#xf4;te d'Ivoire. Children aged 6-12 years were screened between July 14, 2022, and July 31, 2022; children positive for T trichiura on duplicate Kato-Katz smears and with infection intensity of 200 eggs per gram or more were eligible and treated first with albendazole (400 mg) and ivermectin (200 &#x3bc;g/kg) then with oxantel pamoate (20 mg/kg). We constructed a species tree of the Trichuris genus using 12&#x2009;434 orthologous groups. We sequenced individual worms, which were used to confirm the phylogenetic placement and investigate patterns of adaptation through comparative genomic analyses. Finally, we conducted a GWAS to compare albendazole-ivermectin sensitive worms to drug non-sensitive worms. FINDINGS: 670 children were screened, of whom 243 were enrolled and from whom 271 worms were isolated after the first treatment and 827 worms after the second treatment. Sufficient DNA was recovered from 747 worms of which 721 were suitable for further bioinformatic analysis; of these, 179 were albendazole-ivermectin sensitive worms and 542 were drug non-sensitive worms. We present and characterise a new, human-infecting Trichuris species named T incognita n sp, which is morphologically indistinguishable from T trichiura, but forms a distinct phylogenetic clade, closer to Trichuris suis than to the canonical human-infective T trichiura. Comparative genomic analysis of genes suspected to confer resistance to either albendazole or ivermectin in helminths revealed a high number of &#x3b2;-tubulin orthologs, present in the whole population of T incognita n sp, compared with the canonical T trichiura species, but these genes were not associated with a resistant phenotype. The GWAS did not provide conclusive evidence of adaptation to drug pressure within the same species. INTERPRETATION: Our results demonstrate that trichuriasis can be caused by multiple whipworm species, and that differences in response rates might result from species responding differently to drug treatment, rather than from the intraspecies establishment of resistance. This discovery, coupled with the high tolerability of T incognita n sp to albendazole-ivermectin, marks a substantial shift in how we understand and approach whipworm infections. FUNDING: European Research Council.

Trichuris↗

Chromosome level genome assembly and full-length transcriptome of blacktip trevally (Caranx heberi).

Caranx heberi (Bennett, 1830) commonly known as the blacktip trevally belongs to the family Carangidae and is a potential brackishwater aquaculture species. However, the limited genomic resources are hindering the efforts to study its genetic traits and their molecular basis. To bridge this gap, we generated a high-quality reference genome employing multiple sequencing strategies including PacBio Hifi reads (135x), Illumina short reads (150x), and Hi-C chromosome conformation capturing (180x). The high-quality genome assembly consisted of 159 scaffolds summing to 618.71&#x2009;Mb and an N50 value of 26.72&#x2009;Mb. Among these, 24 chromosome level scaffolds covered 97.5% of the total assembly. The genome contained 20.94% of repeat elements and 30,354 protein encoding genes. In addition, full-length transcriptomes were generated using the PacBio IsoSeq approach from seven tissues (gill, kidney, liver, muscle, heart, spleen, and intestine). The comprehensive genomic and transcriptomic resources developed in this study will facilitate the domestication and aquaculture development of C. heberi, as well as support research on its nutritional potential, ecological adaptations, and evolutionary biology.

Animals↗

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis↗

Branch migration inhibition in PCR-amplified DNA: homogeneous mutation detection.

A novel method for detection of any mutation located within a PCR-amplified DNA sequence was demonstrated. The method is based on the inhibition of spontaneous DNA branch migration. Partial duplexes produced by PCR amplification of a test and a reference genomic DNA sample anneal to form four-stranded cruciform structures. Spontaneous DNA branch migration results in dissociation of these structures when the test and reference sequences are identical. Any base substitution, deletion or insertion inhibits branch migration and produces stable cruciform structures. When suitable ligands are attached to the PCR primers, the cruciform structures can be detected by standard immunochemical methods. This approach was tested using several commonly occurring mutations within the human cystic fibrosis gene. New methods for increasing the specificity of PCR amplifications are described that were used for successful mutation analysis.

Cystic Fibrosis↗

Conserved host-exclusive oligonucleotide motifs enriched in pathogenic genes of human oncogenic viruses.

Comparative viral genomics can reveal sequence-level constraints influencing virus-host interactions. Relative minimal absent words (rMAWs) are short oligonucleotide motifs present in viral genomes but completely absent from the host, potentially reflecting selective pressures related to host adaptation and immune evasion. Using the EAGLE algorithm and the GRCh38 human reference genome, we systematically screened for prevalent rMAWs (prMAWs) across six major human oncogenic viruses: Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell leukemia virus type 1 (HTLV-1), and human herpesvirus 8/Kaposi's sarcoma-associated herpesvirus (HHV-8/KSHV). highly conserved 11- and 12-bp prMAWs were identified in EBV, HBV, HTLV-1, and HHV-8/KSHV, with sequence prevalences ranging from 91.5% to 97.9%. Conversely, no short prMAWs were detected in HCV or HPV, likely reflecting differences in genome architecture, mutation rates, and long-term host adaptation to the human host. Importantly, the identified host-exclusive motifs exhibited non-random genomic distribution and were preferentially embedded within viral genes central to replication, persistence, immune modulation, and oncogenesis, including EBNA-1 (EBV), HBx (HBV), Tax-associated regions (HTLV-1), and lytic replication genes of HHV-8/KSHV. Notably, all detected prMAWs were enriched in GC nucleotides and exhibited marked CpG over-representation, suggesting sequence constraints associated with epigenetic regulation and viral persistence. Collectively, these highly conserved, host-exclusive signatures offer promising, candidates for sequence-directed approaches in the diagnosis, monitoring, and investigation of virus-associated cancers.

Humans↗

The genome sequence of the comb-clawed beetle, Prionychus ater (Fabricius, 1775) (Coleoptera: Tenebrionidae).

We present a genome assembly from an individual male Prionychus ater (comb-clawed beetle; Arthropoda; Insecta; Coleoptera; Tenebrionidae). The assembly contains two haplotypes with total lengths of 385.22 megabases and 347.54 megabases. Most of haplotype 1 (97.6%) is scaffolded into 14 chromosomal pseudomolecules, including the X sex chromosome. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 16.26 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain&#x202f;and&#x202f;Ireland.

Prionychus ater; comb-clawed beetle; genome sequen↗

The genome sequence of the Treble Lines, Charanyca trigrammica (Hufnagel, 1766) (Lepidoptera: Noctuidae).

We present a genome assembly from an individual male Charanyca trigrammica (Treble Lines; Arthropoda; Insecta; Lepidoptera; Noctuidae). The assembly contains two haplotypes with total lengths of 546.43 megabases and 546.58 megabases. Most of haplotype 1 (99.97%) is scaffolded into 31 chromosomal pseudomolecules, including the Z sex chromosome. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 15.44 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain&#x202f;and&#x202f;Ireland.

Charanyca trigrammica; Treble Lines; genome sequen↗

The genome sequence of the Eurasian Spoonbill, Platalea leucorodia Linnaeus, 1758 (Pelecaniformes: Threskiornithidae).

We present a genome assembly from an individual female Platalea leucorodia (Eurasian Spoonbill; Chordata; Aves; Pelecaniformes; Threskiornithidae). The assembly contains two haplotypes with total lengths of 1&#x202f;345.83 megabases and 1&#x202f;190.44 megabases. Most of haplotype 1 (95.57%) is scaffolded into 37 chromosomal pseudomolecules, including the W and Z sex chromosomes. Haplotype 2 was assembled to scaffold level. The mitochondrial genome has also been assembled, with a length of 17.15 kilobases. This assembly was generated as part of the Darwin Tree of Life project, which produces reference genomes for eukaryotic species found in Britain&#x202f;and&#x202f;Ireland.

Platalea leucorodia; Eurasian Spoonbill; genome se↗