Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Simultaneous detection of Dialister pneumosintes and Filifactor alocis in endodontic infections by 16S rDNA-directed multiplex PCR.

Dialister pneumosintes and Filifactor alocis have been recently considered as candidate endodontic pathogens. In this study, we devised a 16S rDNA-directed multiplex PCR protocol for simultaneous detection of these two bacterial species in endodontic infections. Samples were taken from infected root canals associated with asymptomatic periradicular lesions as well as from cases of acute periradicular abscesses. DNA extracted from the samples was used as template for simultaneous detection of D. pneumosintes and F. alocis through a multiplex PCR assay. Two fragments of the expected sizes, one specific for D. pneumosintes and the other for F. alocis, were simultaneously amplified from a mixture of reference genomic DNA containing DNA from both species. Clinical samples that were positive for the target species showed a single band of the predicted size for each species. D. pneumosintes was detected by multiplex PCR in 11 samples (7 asymptomatic and 4 abscesses) and F. alocis was identified in 9 cases (6 asymptomatic and 3 abscesses). Six samples (3 asymptomatic and 3 abscesses) shared the two species. Data from the present study confirmed that D. pneumosintes and F. alocis are common members of the microbiota present in primary endodontic infections and thereby may participate in the pathogenesis of periradicular lesions. The proposed multiplex PCR assay is a simple, rapid, and accurate method for the simultaneous detection of these two candidate endodontic pathogens.

Adult↗

Disruption of GAD1 protein architecture by a novel missense variant in a consanguineous family with autosomal recessive intellectual disability.

BACKGROUND: Intellectual disability represents a heterogeneous group of neurodevelopmental disorders marked by significant impairments in intellectual functioning and adaptive behavior. Among the various causes, genetic factors play a major role, with autosomal recessive intellectual disability (ARID) constituting a genetically diverse subgroup. ARID is prevalent in consanguineous families and arises from homozygous mutations that disrupt critical genes involved in brain development and function. OBJECTIVE: This study aimed to identify disease-causing genetic variants responsible for ARID in a consanguineous Pakistani family and to evaluate the structural and functional impact of a novel variant identified in GAD1 through protein modeling. METHODS: A consanguineous family affected with intellectual disability was enrolled. Whole-exome sequencing was performed on an affected individual, followed by bioinformatics analysis including alignment to the GRCh38 reference genome, variant calling, and annotation. Variants were filtered based on rarity, predicted functional impact, and autosomal recessive inheritance pattern. Candidate variants were validated and assessed by Sanger sequencing and segregation analysis. Protein modeling was performed to evaluate the structural impact of the identified variant. RESULTS: A novel homozygous missense variant NM_000817:c.1700G>A;p.Arg567Gln in GAD1 was identified. Segregation analysis confirmed co-segregation of the variant with the affected phenotype. Protein modeling suggested that the variant may disrupt GAD1 enzymatic function involved in gamma-aminobutyric acid synthesis. CONCLUSION: This study emphasizes the significance of genetic investigation in familial cases and the crucial role that GAD1 mutations play in neurodevelopmental disorders with intellectual disability. The results advance the knowledge of molecular causes of ARID and broaden the mutational range.

Pakistani↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype↗

Large-scale production of SAGE libraries from microdissected tissues, flow-sorted cells, and cell lines.

We describe the details of a serial analysis of gene expression (SAGE) library construction and analysis platform that has enabled the generation of >298 high-quality SAGE libraries and >30 million SAGE tags primarily from sub-microgram amounts of total RNA purified from samples acquired by microdissection. Several RNA isolation methods were used to handle the diversity of samples processed, and various measures were applied to minimize ditag PCR carryover contamination. Modifications in the SAGE protocol resulted in improved cloning and DNA sequencing efficiencies. Bioinformatic measures to automatically assess DNA sequencing results were implemented to analyze the integrity of ditag structure, linker or cross-species ditag contamination, and yield of high-quality tags per sequence read. Our analysis of singleton tag errors resulted in a method for correcting such errors to statistically determine tag accuracy. From the libraries generated, we produced an essentially complete mapping of reliable 21-base-pair tags to the mouse reference genome sequence for a meta-library of approximately 5 million tags. Our analyses led us to reject the commonly held notion that duplicate ditags are artifacts. Rather than the usual practice of discarding such tags, we conclude that they should be retained to avoid introducing bias into the results and thereby maintain the quantitative nature of the data, which is a major theoretical advantage of SAGE as a tool for global transcriptional profiling.

Animals↗

Genetic Adaptation to Brackish Water and Spawning Season in European Cisco.

How species adapt to diverse environmental conditions is essential for understanding evolution and the maintenance of biodiversity. The European cisco (Coregonus albula) is a salmonid that occurs in both fresh and brackish water, and this together with the presence of sympatric spring- and autumn-spawning lacustrine populations provides an opportunity for studying the genetics of adaptation in relation to salinity and timing of reproduction. Here, we present a high-quality reference genome of the European cisco based on PacBio HiFi long read sequencing and HiC-directed scaffolding. We generated low-coverage whole-genome sequencing data from 336 individuals across 12 population samples to explore population structure and genetics of ecological adaptation. We found a major subdivision between two groups of populations most likely reflecting colonisation from different glacial refugia. Within the two major groups, we detected further genetic differentiation between spring- and autumn-spawning populations and between populations from freshwater lakes, rivers and brackish water (Bothnian Bay). A genome-wide screen for genetic differentiation among populations identified a set of outlier SNPs strongly correlated with spawning timing and salinity. Several of the genes associated with spawning time, including BHLHE40, TIMELESS and CPT1A, have previously been shown to have a role in circadian rhythm biology. As many as 17 loci were associated with genetic differentiation between populations reproducing in fresh and brackish water. This study provides insights into the genomic basis of ecological adaptation in European cisco with implications for sustainable fishery management.

Animals↗

Multivariate Effects of SNPs on Environmental Streptococcal Mastitis Evaluated With an NGS-Based Association Study Using Targeted Resequencing in the Bovine MHC Region.

Mastitis is an inflammatory reaction caused by bacterial infection of the teat, and a relationship between its onset and cattle major histocompatibility complex (BoLA) region has been reported. However, no comprehensive genetic analysis of mastitis caused by environmental streptococci has been reported. Here, we resequenced the BoLA region using a hybridisation capture target next-generation sequencing (NGS) method to identify disease susceptibility markers mapped to the BoLA region in environmental streptococcal mastitis. This study examined 75 cows with mastitis caused by environmental streptococci selected from 1641 cows with mastitis and 222 healthy cows without mastitis in Japan. Targeted sequences obtained from MiSeq NGS were aligned to the bovine reference genome (ARS-UCD1.2/bosTau9), and 2,920,355 variants were detected within the BoLA region of the 297 Holstein cattle. In an association study using 2264 variants after quality control, the top 20 variants with the lowest P values were selected and assigned to the 18 surrounding candidate genes, and a gene network analysis of these genes resulted in the narrowing down of five candidate genes POU5F1, IER3, GNL1, ABCF1, and PRR3. Multivariate effect analysis of all 6 SNPs associated with these 5 genes revealed that they were significantly correlated with mastitis, indicating that they were useful for classification of mastitis-resistant and mastitis-susceptible cattle. This is the first report to identify SNPs associated with environmental streptococcal mastitis with an NGS-based association study using targeted resequencing in the BoLA region, and understanding host factors may provide important clues for mastitis control.

Animals↗

Harnessing fern stress adaptations: From evolution and ecophysiology to molecular biology.

Ferns are the second most diverse vascular plant lineage after angiosperms and have been a key ecological component of Earth's biodiversity for more than 380 million years. Importantly, ferns are sister to seed plants, providing a critical outgroup for understanding the evolution of seed plant features. Ferns are remarkably resilient to abiotic and biotic stresses due to a long evolutionary history with adaptations to diverse habitats, stresses, and herbivores. As a result, ferns produce a multitude of secondary metabolites with unique bioactivities; these chemicals are potentially linked to the adaptation of ferns to herbivory, various abiotic and biotic stresses, and changing environments. Assembled reference genomes and the identification of key metabolic compounds of multiple ferns have already made significant contributions to human health and well-being. Here, we review the recent scientific advances in fern research, including evolution, stress resistance, metabolites and medicinal utilization, and comparative multi-omics applications. We propose that integrated investigations involving ecological, physiological, and molecular techniques will facilitate the future research translation of fern resources in diverse areas including soil remediation, biopesticides, and medicine. Advances in our understanding of fern molecular biology will provide new insights into the evolution of land plants and promote the utilization of ferns for heightened environmental restoration, crop protection and human health.

Ferns↗

Recombinant adenoviruses with large deletions generated by Cre-mediated excision exhibit different biological properties compared with first-generation vectors in vitro and in vivo.

In vivo gene transfer of recombinant E1-deficient adenoviruses results in early and late viral gene expression that elicits a host immune response, limiting the duration of transgene expression and the use of adenoviruses for gene therapy. The prokaryotic Cre-lox P recombination system was adapted to generate recombinant adenoviruses with extended deletions in the viral genome (referred to here as deleted viruses) in order to minimize expression of immunogenic and/or cytotoxic viral proteins. As an example, an adenovirus with a 25-kb deletion that lacked E1, E2, E3, and late gene expression with viral titers similar to those achieved with first-generation vectors and less than 0.5% contamination with E1-deficient virus was produced. Gene transfer was similar in HeLa cells, mouse hepatoma cells, and primary mouse hepatocytes in vitro and in vivo as determined by measuring reporter gene expression and DNA transfer. However, transgene expression and deleted viral DNA concentrations were not stable and declined to undetectable levels much more rapidly than those found for first-generation vectors. Intravenous administration of deleted vectors in mice resulted in no hepatocellular injury relative to that seen with first-generation vectors. The mechanism for stability of first-generation adenovirus vectors (E1a deleted) appeared to be linked in part to their ability to replicate in transduced cells in vivo and in vitro. Furthermore, the deleted vectors were stabilized in the presence of undeleted first-generation adenovirus vectors. These results have important consequences for the development of these and other nonintegrating vectors for gene therapy.

Adenovirus E1 Proteins↗

Release of virus-like particles from cells infected with poliovirus replicons which express human immunodeficiency virus type 1 Gag.

The effectiveness of attenuated poliovirus vaccines when given orally to induce both systemic and mucosal immune responses against poliovirus has resulted in an effort to develop poliovirus-based vectors to express foreign proteins. We have previously described the construction of poliovirus genomes (referred to as replicons) in which the complete human immunodeficiency virus type 1 (HIV-1) gag gene was substituted for the capsid gene (P1) (D.C. Porter, D.C. Ansardi, and C.D. Morrow, J. Virol. 69:1548-1555, 1995). Infection of cells with encapsidated replicons resulted in the expression of a 55-kDa protein. To further characterize the biological features of the HIV-1 Gag proteins expressed in cells infected with encapsidated replicons, we utilized biochemical analysis and electron microscopy. Expression of the 55-kDa protein in cells infected with encapsidated replicons resulted in myristylation of the Pr55gag protein. The Gag precursor protein was released from infected cells; analysis on sucrose density gradients revealed that the precursor sedimented at a density consistent with that of an HIV-1 virus-like particle. Analysis of replicon-infected cells by electron microscopy demonstrated the presence of condensed structures at the plasma membrane and the release of virus-like particles. These studies demonstrate that poliovirus-based vectors can be used to express foreign proteins which require posttranslational modifications, such as myristylation, and assemble into higher-order structures, providing a foundation for the future use of poliovirus replicons as vaccine vectors.

Amino Acid Sequence↗

Alternative quadruplex real-time PCR reactions for detection and discrimination of Streptococcus pneumoniae serotypes within serogroup 6.

UNLABELLED: Streptococcus pneumoniae causes significant morbidity and mortality worldwide, and serotyping is important to assess the burden of disease that is vaccine preventable. For serotyping, the Centers for Disease Control and Prevention (CDC) use a series of 12 real-time multiplex PCRs (rmPCRs) performed in quadruplex reactions; however, rmPCR reaction 5 (rmPCR-5) for serotypes 6A, 6B, 6C, and 6D often failed at low DNA concentrations. This study investigated the cause of rmPCR-5 failure and provided alternative rmPCRs to resolve this issue. Quadruplex rmPCR target sequences were compared to S. pneumoniae reference genomes. Reactions rmPCR-5 [6ABCD, 6AB, 6BD, and 6CD] and rm-PCR-11 [37, 10F, 11BC, and 18CFBA] were compared to alternative reactions rmPCR-A1 [6ABCD, 10F, 11BC, and 18CFBA] and rmPCR-A2 [37, 6AB, 6BD, and 6CD]. All rmPCRs were tested using 10-fold serial dilutions of DNA from representative serotypes, and analytical specificity was assessed using DNA from other S. pneumoniae serotypes or various streptococci and Gram-positive cocci. Failure of rmPCR-5 was associated with overlapping 6ABCD and 6BD targets. Separation of these targets in the alternative rmPCRs-A1 and rmPCR-A2 allowed sensitive and specific detection and discrimination of serotypes 6A, 6B, 6C, and 6D, without impacting the detection of serotypes 10F, 11BC, 18CFBA, and 37. This study highlights the importance of rigorous author and peer-review to avoid manuscript errors and unintended consequences. By explaining what caused rmPCR-5 failure and proposing alternative reactions rmPCRs-A1 and rmPCR-A2, this study demonstrates the value of scientific collaboration to ensure molecular assays best serve the scientific community. IMPORTANCE: Streptococcus pneumoniae is a bacterium that can cause life-threatening infections like pneumonia and meningitis, leading to millions of deaths worldwide each year. A key feature enabling S. pneumoniae to cause disease is its sugar coating, allowing it to avoid the immune system. These surface sugars are the target of S. pneumoniae vaccines. However, vaccines only protect against some sugars and understanding which ones are on the surface of S. pneumoniae is called "serotyping." The Centers for Disease Control and Prevention (CDC) have protocols that allow us to predict S. pneumoniae serotypes by looking at its DNA. We found errors in the CDC protocols and provided a simple solution to fix them. Ultimately, having accurate serotyping protocols allows us to know how much disease is preventable by vaccine, allows us to monitor how well vaccine are working, and helps develop new vaccines if needed.

Streptococcus pneumoniae↗

ESTIMA, a tool for EST management in a multi-project environment.

BACKGROUND: Single-pass, partial sequencing of complementary DNA (cDNA) libraries generates thousands of chromatograms that are processed into high quality expressed sequence tags (ESTs), and then assembled into contigs representative of putative genes. Usually, to be of value, ESTs and contigs must be associated with meaningful annotations, and made available to end-users. RESULTS: A web application, Expressed Sequence Tag Information Management and Annotation (ESTIMA), has been created to meet the EST annotation and data management requirements of multiple high-throughput EST sequencing projects. It is anchored on individual ESTs and organized around different properties of ESTs including chromatograms, base-calling quality scores, structure of assembled transcripts, and multiple sources of comparison to infer functional annotation, Gene Ontology associations, and cDNA library information. ESTIMA consists of a relational database schema and a set of interactive query interfaces. These are integrated with a suite of web-based tools that allow a user to query and retrieve information. Further, query results are interconnected among the various EST properties. ESTIMA has several unique features. Users may run their own EST processing pipeline, search against preferred reference genomes, and use any clustering and assembly algorithm. The ESTIMA database schema is very flexible and accepts output from any EST processing and assembly pipeline. ESTIMA has been used for the management of EST projects of many species, including honeybee (Apis mellifera), cattle (Bos taurus), songbird (Taeniopygia guttata), corn rootworm (Diabrotica vergifera), catfish (Ictalurus punctatus, Ictalurus furcatus), and apple (Malus x domestica). The entire resource may be downloaded and used as is, or readily adapted to fit the unique needs of other cDNA sequencing projects. CONCLUSIONS: The scripts used to create the ESTIMA interface are freely available to academic users in an archived format from http://titan.biotec.uiuc.edu/ESTIMA/. The entity-relationship (E-R) diagrams and the programs used to generate the Oracle database tables are also available. We have also provided detailed installation instructions and a tutorial at the same website. Presently the chromatograms, EST databases and their annotations have been made available for cattle and honeybee brain EST projects. Non-academic users need to contact the W.M. Keck Center for Functional and Comparative Genomics, University of Illinois at Urbana-Champaign, Urbana, IL, for licensing information.

Animals↗

Characterization and analysis of the full-length transcriptome of Frankliniella occidentalis (Thysanoptera: Thripidae).

BACKGROUND: Frankliniella occidentalis, an insect belonging to the order Thysanoptera, causes severe damage to agricultural and horticultural crops, resulting in significant economic losses worldwide. The development of molecular and sequencing technologies has helped elucidate the molecular mechanisms regulating its growth and development as well as its damaging activity. However, much remains to be explored. To further investigate the molecular complexity of this species, we sequenced the full-length transcriptome of mixed samples obtained from specimens at all developmental stages. RESULTS: Of all transcripts, 89.04% matched with the reference genome; additionally, 29,750 alternative splicing events, 2,342 genes with poly(A) sites, and 153 candidate fusion transcript events were identified, and 4,235 long noncoding RNAs were discovered. CONCLUSIONS: This is the first full-length transcriptome of F. occidentalis reported to date. This study greatly contributes to the understanding of the molecular complexity and diversity of this insect, providing a basis to develop specific molecular targets as well as resources for gene function studies in other insects.

Animals↗

Identification and fine mapping of a locus controlling multi-main-stem trait in Brassica napus.

BACKGROUND: The main stem is a crucial component determining individual plant yield in rapeseed (Brassica napus). However, the genetic and developmental basis underlying the multi-main-stem trait remains largely unclear. RESULTS: In this study, we identified a multi-main-stem mutant, mms1, which exhibited a significantly increased silique number per plant and abnormal shoot apical meristem (SAM) development. Genetic analysis demonstrated that the multi-main-stem trait was controlled by a recessive gene. Using bulked segregant analysis combined with a Brassica napus 50 K SNP array and map-based cloning, the locus was mapped to a 340-kb interval on chromosome A09 of the ZS11 reference genome and was designated BnaA09.MMS1. Candidate gene analysis revealed that BnaA09G0254500ZS, which harbors sequence variations in both the promoter and coding regions and shows significantly increased expression in the mutant, was the most likely candidate gene. In addition, phytohormone analysis revealed reduced auxin accumulation in mutant SAMs, together with transcriptomic changes in genes associated with the CLAVATA3 (CLV3)-WUSCHEL (WUS) feedback loop. CONCLUSIONS: These findings provide an important foundation for elucidating the genetic basis of the multi-main-stem trait and offer a valuable genetic resource for rapeseed improvement.

Brassica napus↗

Transcriptome analysis of the diseased intervertebral disc tissue in patients with spinal tuberculosis.

OBJECTIVE: To investigate the differential expression genes (DEGs) in spinal tuberculosis using transcriptomics, with the aim of identifying novel therapeutic targets and prognostic indicators for the clinical management of spinal tuberculosis. METHODS: Patients who visited the Department of Orthopedics at the Second Hospital, Lanzhou University from January 2021 to May 2023 were enrolled. Based on the inclusion and exclusion criteria, there were 5 patients in the test group and 5 patients in the control group. Total RNA was extracted and paired-end sequencing was conducted on the sequencing platform. After processing the sequencing data with clean reads and annotating the reference genome, FPKM normalization and differential expression analysis were performed. The DEGs and long non-coding RNAs (LncRNAs) were analyzed for Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) enrichment. The cis-regulation of differentially expressed mRNAs (DE mRNAs) by LncRNAs was predicted and analyzed to establish a co-expression network. RESULTS: This study identified 2366 DEGs, with 974 genes significantly upregulated and 1392 genes significantly downregulated. The upregulated genes are associated with cytokine-cytokine receptor interactions, tuberculosis, and TNF-α signaling pathways, primarily enriched in biological processes such as immunity and inflammation. The downregulated genes are related to muscle development, contraction, fungal defense response, and collagen metabolism processes. Analysis of LncRNAs from bone tuberculosis RNA-seq data detected a total of 3652 LncRNAs, with 356 significantly upregulated and 184 significantly downregulated. Further analysis identified 311 significantly different LncRNAs that could cis-regulate 777 target genes, enriched in pathways such as muscle contraction, inflammatory response, and immune response, closely related to bone tuberculosis. There are 51 genes enriched in the immune response pathway regulated by cis-acting LncRNAs. LncRNAs that regulate immune response-related genes, such as upregulated RP11-451G4.2, RP11-701P16.5, AC079767.4, AC017002.1, LINC01094, CTA-384D8.35, and AC092484.1, as well as downregulated RP11-2C24.7, may serve as potential prognostic and therapeutic targets. CONCLUSION: The DE mRNAs and LncRNAs in spinal tuberculosis are both associated with immune regulatory pathways. These pathways promote or inhibit the tuberculosis infection and development at the mechanistic level and play an important role in the process of tuberculosis transferring to bone tissue.

Humans↗

Small serine recombinases are markers for antiphage defense system discovery.

Renewed interest in phage therapy has highlighted a need to understand how bacteria subvert phage infection through antiphage defense systems. Traditionally, strategies to identify antiphage defense systems lack throughput or have limitations for bacterial species where antiphage defense systems are understudied. Herein, we developed a bioinformatic pipeline that uses a small serine recombinase to identify known and unknown antiphage defense systems. Using this approach to query reference genomes and metagenomes, we show that small serine recombinase genes are genetically linked to antiphage defense systems and serve as bait for finding these systems across diverse bacterial phyla. Using co-transcription predictions and statistical analysis of protein domain abundances, we experimentally validated our bioinformatic approach by discovering that KAP P-loop NTPases are fused to putative antiphage domains and reinforce prokaryotic Schlafen proteins as a new class of antiphage defense. Our work shows that small serine recombinases are a reliable genetic marker for the discovery of antiphage defenses across diverse bacterial phyla.

Bacteriophages↗

De novo assembly and authentication of ancient DNA metagenomes with nf-core/mag.

Ancient DNA provides a direct window into the evolutionary processes that have shaped living microbial species today, as well as their now extinct relatives. Advances in both sequencing methods and de novo assembly techniques have not only resulted in a flood of modern metagenomic sequencing data, but they have also allowed palaeogenomicists to retrieve vast amounts of ancient DNA from past microorganisms, including species and strains without modern reference genomes. However, the degraded nature of ancient DNA means that the standard techniques of genome assembly developed for modern DNA are unlikely to perform effectively, unless heavily modified. This hinders the incorporation of ancient data into broader metagenomic studies that would otherwise benefit from having deep time information on the evolution of different microbial species. In this primer and protocol paper, we provide guidance on ways to adapt existing metagenomic de novo assembly processes, including data input, tools, and settings, in order to perform more robustly and effectively on ancient DNA. After assembly, we then further describe how ancient DNA contigs can be identified and validated. The key steps of ancient metagenomic assembly are now integrated in a dedicated ancient DNA mode in the established pipeline nf-core/mag. By introducing support for ancient DNA data in nf-core/mag, we aim to improve the ability of researchers to more regularly integrate de novo assembled ancient microbial data into broader metagenomics studies of microbial ecology and evolution.

DNA, Ancient↗

Methods for single-pair Ascaridia galli genetic crosses.

Ascarid parasites infect a wide range of hosts, causing significant clinical and economic impacts. However, genetic tools for studying ascarid biology remain limited. We optimized genetic crosses using Ascaridia galli , a common ascarid of chickens. Sexually immature larval parasites were recovered from donors, transferred to gelatin capsules, and then given orally to recipients. We successfully established single-pair matings in 32% of crossing attempts. This method to control genetic crosses further establishes the avian model for ascarid research and will enable future studies to create a high-quality reference genome, inbreed anthelmintic resistant and sensitive lines, and investigate host-pathogen interactions.

Journal Article↗