Search PubMedSearch

SEARCH · Search PubMed

Results for “Low-pass sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

10 recordsLinked to original sources

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals

Clinical and genetic analysis of a family with 16p11.2 microduplication syndrome and variable multisystem manifestations.

16p11.2 microduplication syndrome (OMIM #614671) is a pathogenic recurrent copy-number gain at the 16p11.2 locus and is associated with variable expressivity across neurodevelopmental, growth, and medical phenotypes. Gastrointestinal symptoms have been reported in carrier cohorts, but detailed documentation of gastrointestinal motility and neuromuscular findings remains limited. We performed clinical and genetic analyses in a multigenerational family in which the proband (III1) presented with limb muscle pain, exercise intolerance, and chronic gastrointestinal symptoms. Next-generation sequencing (NGS), low-pass whole-genome sequencing (lpWGS)-based CNV analysis, Sanger sequencing, and qPCR validation identified a 0.8 Mb microduplication at 16p11.2 (BP4-BP5), involving 44 genes including TBX6, inherited from the mother (II2). The proband's clinical manifestations included developmental delay, pointed chin, low body mass index, gastrointestinal dysfunction (chronic abdominal pain, diarrhea, esophageal motility disorder, and rectal prolapse), forward-leaning gait, mild scoliosis, and limb muscle atrophy with inflammatory muscle involvement. Four family members (II2, III1, III2, and III4) carried the microduplication, but their available clinical features varied in severity and system involvement. The proband's twin brother (III2) had left ear deafness and epilepsy, individual II2 had blindness from cone-rod dystrophy, and III4 showed more pronounced scoliosis. This family provides a detailed clinical and genetic description of 16p11.2 microduplication carriers with prominent gastrointestinal motility and neuromuscular manifestations, thereby enriching the clinical characterization of this recurrent CNV and supporting substantial intrafamilial phenotypic heterogeneity.

16p11.2 microduplication syndrome

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99×) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including ~ 17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8 ± 8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (π = 0.00267), followed by lowland (π = 0.00233), whereas highland chickens showed the lowest diversity (π = 0.00203) and elevated genomic inbreeding (FROH and FHOM ≈ 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Identification of novel HUWE1 variants in Turner-type X-linked intellectual disability.

OBJECTIVE: To characterize the clinical phenotypes and identify the genetic etiology in four unrelated families affected by Turner-type X-linked intellectual disability (XLID). METHODS: Peripheral blood samples were collected from four probands and their parents. Genomic DNA was extracted, and a comprehensive genetic analysis was performed using trio-based Whole Exome Sequencing (WES) combined with low-pass Copy Number Variation sequencing (CNV-seq). Candidate variants were subsequently validated via Sanger sequencing. RESULTS: Genetic analysis identified distinct variants in the HUWE1 across the four families. Specifically, four distinct HUWE1 variants were identified across the families: a hemizygous c.10034 > T (p.Lys3345Met) in Family 1; a heterozygous c.9209G > A (p.Arg3070His) in Family 2; a heterozygous c.12688T > C (p.Phe4230Leu) in Family 3; and a hemizygous c.9070G > A (p.Ala3024Thr) in Family 4. In accordance with ACMG guidelines, the novel variants in Families 1, 3, and 4 were classified as "Likely Pathogenic" (PS2 + PM2_Supporting + PP2 + PP3). In contrast, the previously reported variant in Family 2 was categorized as "Pathogenic" based on the criteria PS2 + PM2_Supporting + PM5 + PP2 + PP3_Moderate. All probands were clinically diagnosed with Turner-type XLID. CONCLUSIONS: This study expands the pathogenic variant spectrum of HUWE1 and provides novel molecular evidence for the clinical diagnosis of Turner-type XLID. These findings are of significant value for genetic counseling, carrier screening, and prenatal diagnosis for the affected families.

Humans

Genomic Analysis of Circulating Tumor Cells at the Single-Cell Level.

Circulating tumor cells (CTCs) have a great potential for noninvasive diagnosis and real-time monitoring of cancer. A comprehensive evaluation of four whole genome amplification (WGA)/next-generation sequencing workflows for genomic analysis of single CTCs, including PCR-based (GenomePlex and Ampli1), multiple displacement amplification (Repli-g), and hybrid PCR- and multiple displacement amplification-based [multiple annealing and loop-based amplification cycling (MALBAC)] is reported herein. To demonstrate clinical utilities, copy number variations (CNVs) in single CTCs isolated from four patients with squamous non-small-cell lung cancer were profiled. Results indicate that MALBAC and Repli-g WGA have significantly broader genomic coverage compared with GenomePlex and Ampli1. Furthermore, MALBAC coupled with low-pass whole genome sequencing has better coverage breadth, uniformity, and reproducibility and is superior to Repli-g for genome-wide CNV profiling and detecting focal oncogenic amplifications. For mutation analysis, none of the WGA methods were found to achieve sufficient sensitivity and specificity by whole exome sequencing. Finally, profiling of single CTCs from patients with non-small-cell lung cancer revealed potentially clinically relevant CNVs. In conclusion, MALBAC WGA coupled with low-pass whole genome sequencing is a robust workflow for genome-wide CNV profiling at single-cell level and has great potential to be applied in clinical investigations. Nevertheless, data suggest that none of the evaluated single-cell sequencing workflows can reach sufficient sensitivity or specificity for mutation detection required for clinical applications.

Carcinoma, Non-Small-Cell Lung

Urinary Small Extracellular Vesicle DNA as a Biomarker for the Non-Invasive Diagnosis of Bladder Cancer.

Existing diagnostic technologies for bladder cancer (BC) suffer from low sensitivity, low specificity, or a lack of validation. Therefore, validated, non-invasive diagnostic biomarkers with high sensitivity and specificity for early detection of BC are needed to complement and improve upon the limitations of existing diagnostic methods. We used low-pass whole genome sequencing (LP-WGS) technology to detect copy number variations (CNVs) in small extracellular vesicle (sEV) DNA isolated from urine samples of patients. Based on these results, we constructed and validated a diagnostic model to differentiate between benign and malignant bladder lesions. We conducted a receiver operating characteristic analysis and calculated the area under the curve (AUC) to evaluate the performance of the diagnostic model. The urine sEV-DNA LP-WGS data revealed CNV differences between benign and malignant samples. The diagnostic model achieved an AUC of 0.953, a sensitivity of 86.7%, and a specificity of 100% in the training cohort and an AUC of 0.985, a sensitivity of 90%, and a specificity of 100% in the validation cohort. Even at the lowest coverage depth of 0.01X, the performance of the diagnostic model remained relatively robust. Notably, the performance of this diagnostic model surpassed that of the biomarker neuron-specific enolase (sensitivity: 85.7% vs. 64.3%; specificity: 100% vs. 87.5%) and urinary cytology (sensitivity: 100% vs. 66.7%; specificity: 100% vs. 94.1%). Our study demonstrates that urine sEV-DNA exhibits high discriminatory power in distinguishing between benign and malignant bladder lesions, making it a promising tool for auxiliary diagnosis of BC.

Humans

A cfDNA fragmentomics classifier for noninvasive differentiation of benign and malignant renal masses.

Noninvasive differentiation of malignant and benign renal masses remains a major clinical challenge, particularly for radiologically indeterminate lesions. Here, we developed and validated a plasma cell-free DNA (cfDNA) fragmentomics-based machine learning classifier for renal mass characterization. The model was trained on 331 participants (171 cancer, 160 benign) and independently validated on 144 participants (73 cancer, 71 benign). Three cfDNA fragmentation features, including copy number variation (CNV), fragmentation-based methylation (FRAGMA), and nucleosome footprint (NF), derived from low-pass whole-genome sequencing, were integrated into an ensemble framework. The model achieved strong discriminative performance, with area under the curve (AUC) values of 0.956 in the training cohort and 0.946 in the validation cohort, outperforming individual feature-based models. At a predefined operating threshold corresponding to 90% sensitivity, specificity reached 0.90 and 0.87, respectively. Notably, most cancer samples exhibited low tumor fraction (TF&#x2009;<&#x2009;3%), yet the model maintained robust performance in low-TF samples (AUCs: 0.952 and 0.941, respectively). Performance remained consistent across tumor stage, grade, and histological subtypes. The classifier also demonstrated potential clinical utility in diagnostically challenging settings, including lipid-poor angiomyolipoma and oncocytoma, with 12 of 13 oncocytoma samples correctly classified in an independent cohort. In addition, the model correctly identified 85.3% of benign masses&#x2009;>&#x2009;4&#xa0;cm, for which surgical intervention is more commonly considered, and 84.6% of malignant tumors&#x2009;&#x2264;&#x2009;4&#xa0;cm, for which management can be challenging. Collectively, these findings support cfDNA fragmentomics as a promising noninvasive liquid biopsy approach for renal mass evaluation and clinical decision-making.

Humans

Development of a low-coverage whole genome sequencing screen for apomixis using a diverse set of Malus germplasm.

In the past decade, plant biologists have made several major discoveries pertaining to the genetic basis of apomixis (clonal propagation by seed) that have shown promise in preserving high-value hybrid rice and sorghum genotypes. This progress was made possible by foundational gene discovery efforts in model species and natural apomicts, but pleiotropic obstacles still limit its broad agricultural adoption, especially in eudicots. Thus, it follows that investigations of novel apomicts should lead to the development of new molecular tools for plant breeding. The two most common ways to identify clonal seed production are flow-cytometry seed screens and genome sequencing to compare the DNA sequences of the maternal parent and progeny, traditionally using low-throughput markers. While flow-cytometry has been the dominant method for more than two decades, it provides indirect information on the genetics of a resulting embryo and can be ineffective in certain species. Here we developed a method using short-read whole-genome sequencing at moderately low coverage (averaging 3X and 6X) to screen diverse Malus genotypes maintained in a USDA germplasm collection for clonal seed production. In total, we sequenced 55 genotypes, 1,216 of their embryos, and identified 17 previously undescribed apomictic genotypes. Several more were detected with the flow cytometry seed screen, which helped resolve certain types of reproduction and sources of noise in low-coverage datasets. This low-pass screening-by-sequencing method is a relatively low-cost, rapid method for detecting apomictic genotypes in diverse plant germplasm and when used thoughtfully in conjunction with flow cytometry, provides a new way to visualize the genetic outcomes of sexual and asexual reproduction in plants.

Apomixis