Search PubMedSearch

SEARCH · Search PubMed

Results for “Variant calling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software

Targeted long-read genomic and epigenomic profiling enhances timely comprehensive variant discovery in hypotonia and muscle weakness.

BACKGROUND: Identifying the genetic basis of hypotonia and muscle weakness is critical for patient management and family counseling. However, diagnosis is often hindered by diverse genomic alterations, including repeat expansions, structural variants (SVs), and methylation defects. Standard-of-care testing, largely based on short-read sequencing, is limited in its ability to detect this heterogeneous variation landscape, leaving many patients undiagnosed or requiring lengthy sequential testing. Long-read sequencing represents a promising solution. However, its application as a first-tier diagnostic assay for hypotonia remains unexplored. METHODS: We retrospectively analyzed 227 patients with hypotonia to assess diagnostic yield, time-to-diagnosis, and costs associated with standard-of-care testing. A long-read whole-genome sequencing (LR-WGS) workflow with targeted analysis of hypotonia-associated genes was developed to detect and prioritize pathogenic SNVs, SVs, and CNVs, repeat expansions, and methylation changes at key disease loci. The workflow was validated in a reference-positive cohort with known diagnoses (n = 15) and applied to an unsolved cohort (n = 14). Variant interpretation followed ACMG guidelines and was confirmed with orthogonal methods. RESULTS: Standard-of-care testing achieved a diagnostic yield of 42% with an average time-to-diagnosis of 68.7 days; however, 30% of diagnosed patients experienced significant delays (average 169 days) due to sequential testing. The LR-WGS based approach identified all known pathogenic variants in the positive cohort, including SMN1 deletions, methylation defects at 15q11.2/Prader-Willi locus, FMR1 repeat expansions, and sequence and copy-number variants in > 100 genes underlying myopathies and muscular dystrophies. The targeted long-read pipeline reduced prioritized variant calls by 97.9-99.9% and, in the unsolved cohort, yielded one definitive diagnosis (de novo COL6A3 deletion) and one possible diagnosis (aberrant methylation and copy number at POMK), for an additional 14% yield. Among patients diagnosed after sequential testing (n = 29), LR-WGS is expected to reduce time-to-diagnosis by ~ 85% and decrease cumulative diagnostic delays, with projected healthcare cost savings of $396,000-439,000. Across the entire 227 patient cohort, LR-WGS is anticipated to reduce testing costs by 6.5%, yielding an average savings of $105 per patient. CONCLUSIONS: LR-WGS enables comprehensive discovery of genomic and epigenomic variants in hypotonia and muscle weakness, improving diagnostic yield, shortening diagnostic timelines, and reducing costs compared with current standard-of-care testing.

Humans

Characterization of variant subclasses of cell lines derived from small cell lung cancer having distinctive biochemical, morphological, and growth properties.

We have described the establishment and biochemical characterization of 50 small cell lung carcinoma (SCLC) cell lines. Further analysis of these data, combined with studies of morphology and growth characteristics, indicates that 35 (70%) of the lines retained typical morphology (SCLC, intermediate subtype), growth characteristics (growth as tightly packed floating cellular aggregates, long doubling times and low colony-forming efficiencies), and biochemical profile (presence of L-dopa decarboxylase, bombesin-like immunoreactivity, neuron-specific enolase, and high concentrations of brain isoenzyme of creatine kinase). They are referred to as classic SCLC lines. The remaining 15 (30%) lines had discordant expression of the biochemical markers; they retained high concentrations of brain isozyme of creatine kinase, but had significantly lower concentrations of neuron-specific enolase and lacked L-dopa decarboxylase and bombesin-like immunoreactivity. These cell lines are called variants. SCLC variant lines could further be divided into (a) biochemical variant lines having variant biochemical profile but retaining typical SCLC morphology and growth characteristics; and (b) morphological variant (SCLC-MV) lines having variant biochemical profile, altered morphology (features of large cell undifferentiated carcinoma) and altered growth characteristics (growth as loosely attached floating aggregates, relatively short doubling times and cloning efficiencies). Fifty-five clones derived from the three SCLC subclasses retained their parental phenotypes. In SCLC-MV lines there was a near constant relationship between variant morphology, altered growth characteristics and amplification of the c-myc oncogene; classic SCLC and biochemical variant SCLC lines were not amplified. Variant morphologies frequently are present in SCLC tumors at autopsy, and most SCLC-MV lines reflect changes that had occurred in the tumors from which they were derived. Because SCLC-MV tumors behave more virulently in the patient and are radioresistant in vitro, these findings are of considerable biological and clinical interest.

Bombesin

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Exploring genetic adaptation and microbial dynamics in engineered anaerobic ecosystems via strain-level metagenomics.

Genetic heterogeneity exists within all microbial populations, with sympatric cells of the same species often exhibiting single-nucleotide variations that influence phenotypic traits, including metabolic efficiency. However, the evolutionary dynamics of these strain-level differences in response to environmental stress remain poorly understood. Here, we present a first-of-its-kind study tracking the adaptive evolution of an anaerobic, carbon-fixing microbiota under a controlled engineered ecosystem focused on carbon dioxide bioconversion into methane. Leveraging strain-resolved metagenomics with an ad hoc variant calling and phasing approach, we mapped mutation trajectories and observed that the two dominant Methanothermobacter species maintained distinct sweeping haplotypes over time, most likely due to niche-specific metabolic roles. By combining population genetic statistics and peptide reconstruction, mer and mcrB genes emerged as potential drivers of archaeal strain-level competition. These findings pave the way for targeted engineering of microbial communities to enhance bioconversion efficiency, with significant implications for sustainable energy and carbon management in anaerobic systems.

Metagenomics

Analysis of a repeat-containing family of Giardia lamblia variant-specific surface protein genes: diversity through gene duplication and divergence.

Giardia lamblia trophozoites express on their surfaces one of a set of cysteine-rich antigenically variant proteins, called variant-specific surface proteins, which comprise the majority of proteins detected by surface labeling. While these VSP proteins may be immunodominant proteins important in the host immune response to G. lamblia, the ability to switch expression from one VSP to another may provide a means for the trophozoites to avoid the host immune response. The first VSP characterized, VSPA6 (from the A6 clone of the WB isolate, originally termed CRP170), contains 18-23 copies of a 65 amino acid repeat. We have now used the repeat as a probe to isolate from a WBA6 genomic library two genes related to vspA6 (called vspA6-S1, vspA6-S2). Sequence analysis of the vspA6-S1 gene revealed nearly two complete copies of the 195 bp repeat and substantial nucleotide and translated amino acid similarity in the coding regions 5' and 3' to the repeats. The vspA6-S2 gene, while still related, showed greater divergence from vspA6 than vspA6-S1 in the nonrepeat coding region and contained nearly four copies of a 201 bp repeat that was 75% identical to the 195 bp vspA6 repeat. These results suggest that gene duplication followed by divergence has played a key role in the generation of the vsp gene repertoire.

Amino Acid Sequence

Genomic Characterization of ETV6::RUNX1-Positive Childhood B-ALL in a Chinese Cohort: Novel Fusion Partners, Co-Occurring Mutations, and Risk-Stratifying Biomarkers.

BACKGROUND: ETV6::RUNX1 is the most common genetic abnormality in pediatric B-cell acute lymphoblastic leukemia (ALL; ∼25%), yet the comprehensive genetic architecture and molecular predictors of intermediate-risk (IR) stratification remain incompletely characterized. METHODS: We performed whole-transcriptome sequencing (Illumina NovaSeq 6000, rRNA depletion, 41.70 Gb/sample) on bone marrow samples from 93 pediatric ETV6::RUNX1-positive B-ALL patients. Bioinformatics analysis included STAR alignment, MuTect2 variant calling, FusionCatcher fusion detection, and VEP annotation. The Jaccard index with permutation testing assessed mutation co-occurrence; logistic regression identified independent predictors of IR classification. RESULTS: Beyond ETV6::RUNX1, we identified 51 distinct fusion genes across the cohort, including the reciprocal RUNX1-ETV6 (73.1%), chr8::KLF1210 (38.7%), and KLF12-chr8 (34.4%). Somatic mutations in 249 genes were detected; the most frequent were KIAA1715 (17.2%), KRAS (11.8%), and NSD2 (10.8%). Network analysis revealed significant chromatin modifier co-occurrence (KIAA1715-KMT2C: J = 0.136, p = 0.015) and KRAS-NRAS mutual exclusivity (J = 0.000, p = 0.042). PTCH1 (OR = 3.50, 95% CI 0.21-58.49, p = 0.41) and GNB1 (OR = 6.5, 95% CI 1.2-34.8, p = 0.029) mutations independently predicted IR classification. chr8::KLF1210 fusion correlated with higher Day-19 MRD levels (p = 0.038). CONCLUSIONS: GNB1 mutation represents a novel independent predictor of IR stratification in ETV6::RUNX1-positive B-ALL. The chromatin modifier co-occurrence module and extensive fusion architecture reveal biological heterogeneity within this favorable-risk subtype, with potential implications for risk-adapted therapeutic strategies.

B‐ALL

Pi subtyping by isoelectric focusing: further genetic studies and application to paternity examinations.

Genetic variation of the protease inhibitor (Pi) alpha 1-antitrypsin was analyzed by isoelectric focusing on polyacrylamide gels in a sample of 347 unrelated individuals from Southern Germany. Six common subtypes of PiM were observed as well as the relatively frequent variants PiS and PiZ and the rare variants PiT, Pi less than L, PiL, PiI and PiF. Also, a variant called PiZ1 was found. The frequency of alleles in this sample was PiM1 = 0.6917, PiM2 - 0.1686, PiM3 = 0.0865, PiS = 0.0230, PiZ = 0.0187, and Pi* = 0.0115. In 82 families the distribution of Pi types was in agreement with an autosomal codominant mode of inheritance. The application of Pi classification in cases of disputed paternity is discussed.

Adult

GD (--) Aachen, a new variant of deficient glucose-6-phosphate dehydrogenase. Clinical, genetic, biochemical aspects.

A deficient G-6PD variant was discovered in 4 males of one family from northwestern Germany. Five generations of this family could be studied. The deficient G-6PD was a new variant, called "Gd (--) Aachen". Its main characteristics are the following: severe enzyme deficiency in erythrocytes (3% of normal), contrasting with an almost normal activity in leukocytes; normal molecular specific activity (i.e., normal ratio enzyme activity/cross-reacting material); slow mobility in starch gel electrophoresis (92-94% of normal); increased Michaelis constant for glucoes-6-phosphate (60-70 muM) and NADP+ (20-25 muM); decreased inhibition constant by NADPH with respect to NADP+ (7 muM); increased inhibition by ATP; normal utilization of the substrate analogues; slightly biphasic pH curve; thermal instability, and normal activation energy of the enzymatic reaction. The relationships between the hematologic disorders (severe and frequent hemolytic crises) and the unfavorable kinetic modifications are discussed.

Adult

A transcriptome-wide approach for rapid pathotype discrimination of Puccinia striiformis f. sp. tritici in north-western India.

Stripe rust of wheat caused by Puccinia striiformis f. sp. tritici (Pst) remains a major constraint to wheat production in India due to the rapid evolution and frequent emergence of virulent pathotypes. Rapid and reliable discrimination of Pst pathotypes is essential for effective resistance deployment and surveillance. In the present study, transcriptome-wide simple sequence repeats (SSRs) and single nucleotide polymorphisms (SNPs) were exploited to develop and validate molecular markers for pathotype-specific detection of Pst pathotypes prevalent in North India (110S119, 238S119, 46S119, 110S84 and 78S84). Microsatellite mining from 6103 core orthologous clusters comprising 51,127 transcripts mined 14,634 SSR loci, from which 93 primer pairs were synthesized. However, only three SSR markers exhibited polymorphism indicating limited discrimination potential of expressed sequence-derived (EST) SSRs for pathotype differentiation. In contrast, SNP discovery through stringent variant calling and filtration yielded 186 pathotype-specific homokaryotic SNPs, of which 56 high-confidence loci were selected for Kompetitive Allele-Specific PCR (KASP) assay development. A total of 48 KASP markers were synthesized and 14 demonstrated clear pathotype- or cluster-specific polymorphism representing substantially higher resolution than SSR markers. The high SNP-to-KASP conversion efficiency (~ 95%) and reproducible fluorescence-based clustering emphasize the robustness of KASP assay. Comparative evaluation revealed that SNP-based KASP markers provide superior discriminatory capacity for closely related Pst pathotypes and represent a promising complementary molecular approach for rapid identification of predominant Indian Pst pathotypes. The validated marker panel developed in this study can complement conventional virulence phenotyping and field pathogenomics approaches for surveillance of currently known pathotypes, while continued refinement may accommodate future changes in pathogen populations.

India

Inhibition of antibody-dependent allergic autocytotoxicity in rheumatoid arthritis by OM-89.

Rheumatoid arthritis (RA) is a disease of multiple etiologies and clinical evidence suggests that a separate variant called "allergic arthritis" induced by food antigens could exist. A missing link in the confirmation of such an observation is a relative lack of a reliable in vitro assay which can confirm the in vivo oral ingestion challenge. Therefore, white blood cells (WBC) from 33 rheumatoid arthritis patients were separated and their disintegration was measured in the presence of specific IgE RAST positive sera and gluten-gliadin antigens. This assay was called the antibody-dependent allergic autocytotoxicity (ACT) test which represents an equivalent of an oral ingestion challenge with food antigens. Control WBC expressed 10-45% disintegration as compared to 75-95% in RA. Preincubation of WBC with OM-89 (immunomodulating fractions of Escherichia coli, OM Laboratories Ltd, Geneva, Switzerland) inhibited significantly antibody-dependent ACT in a dose-related manner (P less than 0.001) in our patients.

Adjuvants, Immunologic

Fluorescent Y-chromosomes in hairs and blood stains.

Y-chromosome detection by way of fluorescence microscopy in biological materials has made sex determination possible in various areas of investigation. The present report describes the results of sex determination on hairs and blood stains. Significant differences were found between the Y-body count for female and male materials. In blind trials it was demonstrated that a reliable sex determination of hairs was possible for at least 27 weeks and of blood stains on cotton cloth and glass for 6 weeks. There were no false positive findings, but there was one male with a "female" blood smear count, who revealed an abnormally small fluorescent region on his Y-chromosome. The existence of such variants calls for caution when evaluating a low count.

Blood Stains

Clinical applications and methodological developments of the RARE technique.

The RARE technique is an extremely useful tool for clinical diagnosis, since it delivers images comparable to those produced by X-ray myelography and X-ray urography. Contrary to these methods, RARE does not use the application of contrast agents. A low-flip-angle variant called FLARE (fast low angle refocused echo imaging) makes this technique accessible for high-field systems.

Humans

Whole genome sequence data on Ethiopian key sorghum landraces and founder lines.

Sorghum (Sorghum bicolor (L.) Moench) is the fifth most important cereal globally. Its genetic diversity is key to improving yield stability, stress tolerance, and adaptation to different environments. Ethiopia is one of the centers for the crop's origin, diversity, and use in both human food and livestock feed. However, genomic data on Ethiopian sorghum remain limited, especially for landraces preferred by local farmers. This dataset consists of whole-genome sequencing data for 188 Ethiopian sorghum accessions, including founder lines and important landraces from major agroecological zones. Sequencing was performed using the Complete Genomics DNBSEQ-T7 platform, generating high-coverage whole-genome data (20 &#xd7; coverage). On average, each accession produced 64.7 million reads. Reads were aligned to the Sorghum bicolor NCBIv3 reference genome and variants called using GATK HaplotypeCaller with joint genotyping (GATK v4.6.1.0). Hard-filtering followed GATK best-practice thresholds (QD <2.0, FS> 60.0, MQ <40.0, MQRankSum <-12.5, ReadPosRankSum <-8.0), retaining biallelic SNPs with mean depth 10-50&#xd7;, missingness &#x2264;20%, and MAF &#x2265;0.05, yielding 6095,752 high-confidence SNPs across 185 accessions. Both the raw FASTQ files and processed VCF files are publicly available to support studies of sorghum genetic diversity, population structure, selection, and the genetic basis of important traits.

Adaptation

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans

Helical structures of poly(D-L-peptides). A conformational energy analysis.

Conformational energy calculations are reported for a number of possible helical structures of poly(D-L-peptides): the alpha helix, two single-stranded piDL, and five double-stranded pipiDL helices. For a poly(D-alanine-L-alanine) sequence, the energies of the various helices are found to differ by less than 1 kcal/(mol residue). For some helices (especially the piDL ones) two structural variants are predicted. These variants, called "goniomers", are characterized by reversed sequences of conformational angles but have the same screw sense and similar helical parameters. A biological implication of these goniomers is suggested, and their usefulness as a critical test for energy calculations is considered.

Alanine

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals

Enhanced reaction with Vicia graminea lectin and exposed terminal N-acetyl-D-glucosaminyl residues on a sample of human red cells with Hb M-Hyde Park.

A sample of polyagglutinable red cells was obtained from a healthy individual (group O, N) possessing a hemoglobin (Hb) variant called Hb M-Hyde Park. The sialic acid content of the individual's red cells is 90 percent of normal, and his cells are agglutinated by monoclonal but not lectin anti-Tn, a panel of lectins specific for N-acetylgalactosamine (or galactose), and N-acetylglucosamine. Enhanced agglutination reactions were obtained with Vicia graminea, Ulex europaeus, and human anti-I and -i. Using various enzyme treatments and different methods of labeling cell surface components, two defective cell membrane sites have been identified: one associated with the O-linked oligosaccharides on sialoglycoproteins and the other associated with exposed N-acetylglucosaminyl residues located on membrane components of apparent molecular weights 88,000 to 130,000 and 46,000 to 73,000 (probably the Band 3 and Band 4.5 regions, respectively).

Acetylglucosamine