Search PubMedSearch

SEARCH · Search PubMed

Results for “Variant calling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

First nationwide full-genome characterisation of human-derived Andes virus in Chile: a retrospective genomic epidemiology study.

BACKGROUND: Andes virus (ANDV) is the only hantavirus known to transmit between humans and causes hantavirus cardiopulmonary syndrome in Chile and Argentina. In Chile, ANDV genomic diversity remains incompletely characterised. This study aimed to characterise the genetic diversity, geographical structure, and molecular signatures of ANDV using human clinical samples collected over a 13-year period (2011-24). METHODS: We conducted a retrospective genomic epidemiology study of ANDV infections in Chile. Clinical samples from patients with confirmed ANDV, collected between March 9, 2011, and June 27, 2024, were analysed and sequenced. Clinical and epidemiological data were obtained from diagnostic laboratories and surveillance programmes. Consensus sequences for the S, M, and L segments were generated, and genetic clustering and divergence were assessed using phylogenetic inference and variant calling. FINDINGS: We analysed clinical samples from 58 infected individuals and identified two major genomic variants of ANDV with distinct geographical distributions, defined by regionally structured patterns of nucleotide and amino acid substitutions across the S, M, and L segments: ANDV Chi-North (central Chile) and ANDV-South (southern Chile). No consistent clustering by clinical severity was observed, and no recurrent non-synonymous substitutions were uniquely associated with severe disease. Substitutions previously associated with person-to-person transmission in outbreaks in Argentina were not consistently observed in Chilean sequences, including in four person-to-person transmission cases. Although some substitutions described in ANDV-like viruses were present in the Chi-North lineage, this lineage remained phylogenetically distinct and geographically restricted to central Chile. INTERPRETATION: To our knowledge, this study provides the first nationwide genomic characterisation of human-derived ANDV in Chile. The identification of geographically structured variants indicates that ANDV diversity in Chile is driven by regional diversification rather than clinical outcome. The absence of consistent amino acid signatures associated with disease severity or person-to-person transmission suggests that these phenotypes are unlikely to be explained by viral genetic variation alone. These findings refine current understanding of ANDV evolution and highlight the need for continued integrated genomic surveillance in endemic regions. FUNDING: Agencia Nacional de Investigación y Desarrollo de Chile and National Institutes of Health.

Humans

Molecular basis of the activation of the tumorigenic potential of Gag-insulin receptor chimeras.

A previous study showed that the human insulin receptor (IR) could be activated by insertion of a 3' portion of the cDNA encoding the beta subunit into a retrovirus genome to form a Gag-IR fusion protein. While capable of transforming cells in culture, this IR cDNA-containing virus, called UIR, was not able to induce tumors in animals. Subsequently, we isolated a spontaneous sarcomagenic variant called UIR19t from the parental UIR. UIR19t was molecularly cloned, sequenced, and found to harbor two mutations. A 44-amino acid deletion immediately upstream from the transmembrane domain of the Gag-IR fusion protein removes all the extracellular sequence of the IR remaining in the original UIR construct. In addition, a single nucleotide deletion at the 3' end results in truncation and replacement of the carboxyl-terminal 12 amino acids by 4 new amino acids. The specific kinase activity of UIR19t is 4- to 5-fold higher than that of the parental UIR. However, no new cellular substrates were detected in UIR19t-transformed cells as compared to UIR cells. Viruses containing either the 5' or the 3' deletion mutation were constructed and assessed for their biological function. Our data indicate that the 5' deletion alone is sufficient to confer tumorigenic ability. We conclude that sequence immediately upstream from the transmembrane domain imposes a negative effect on the transforming and tumorigenic potential of the Gag-IR fusion protein.

Amino Acid Sequence

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA

Characterization of in vitro immunoselected variants from a highly metastatic murine tumor for alterations in malignant behavior in vivo.

A new Ly-6.2- antigen-loss variant (called L61 -M1) of the highly metastatic DBA/2 mouse (Ly-6.2+) MDAY-D2 tumor has been obtained by means of a monoclonal anti-Ly-6.2 antibody in an in vitro immunoselection technique. Whereas L61 -M1 grew poorly when inoculated subcutaneously into the syngeneic host, it grew and metastasized in a similar way to the parental MDAY-D2 tumor when inoculated into immunosuppressed, athymic nude mice. L61 -M1 as well as another Ly-6.2- variant of the same MDAY-D2 tumor (called L61 ) which is poorly metastatic in the syngeneic host salvaged exogenous fucose into glycoproteins and glycolipids at rates 5.5 and 7.8 times that of the parental MDAY-D2 line. In contrast, the Ly-6.2- variants exhibited a 50-70% decrease in the incorporation of exogenous mannose into glycoproteins and glycolipids. L61 -M1 and L61 also exhibited alterations in the structures of the oligosaccharide moieties linked to the cell surface glycoproteins and/or glycolipids. Thus, the in vitro immunoselection technique can be used to obtain a panel of variants with stable phenotypic alterations in their growth and metastatic capacities. Such mutants may, like previously described lectin-resistant mutants, be useful in studying the contribution of cell surface glycoproteins and glycolipids to tumorigenicity and metastasis.

Animals

Promoter-selective activation domains in Oct-1 and Oct-2 direct differential activation of an snRNA and mRNA promoter.

The promoter specificity of transcriptional activators is generally thought to be conferred by the specificity of the DNA-binding domain, which brings the activation domain to the appropriate promoter sequence. We show here, however, that Oct-1 and Oct-2 can differentially activate transcription not through DNA binding specificity but instead through the use of promoter-selective activation domains. These distinct activation domains lead to stimulation of the U2 small nuclear RNA promoter by Oct-1 and an mRNA promoter by Oct-2. An Oct-2 variant, called Oct-2B, differs from Oct-2 by an Oct-1-related C-terminal extension that results from alternative splicing. This variant gains the ability to activate the U2 small nuclear RNA promoter. Thus, the promoter selectivity of a transcriptional activator can be changed, in this case by alternative splicing, without affecting its DNA binding specificity.

DNA-Binding Proteins

3T3 variants unable to bind epidermal growth factor cannot complement in co-culture.

A Swiss-Webster 3T3 variant, called 3T3-ENR7, unable to divide in response to epidermal growth factor was isolated by the mitogen-colchicine selection technique. Like the other EGF non-proliferative variants (3T3-NR6 and 3T3-TNR2), 3T3-ENR7 was unable to bind 125I-EGF. Pairwise co-culture of the three independently isolated EGF non-responsive variants did not restore mitogenic responsiveness to EGF.

Animals

An examination of tumor antigen loss in spontaneous metastases.

Metastases arising from a subcutaneous injection of the DBA/2 tumor, MDAY-D2, as well as four drug-resistant variants (either wheat germ agglutinin-resistant, ouabain-resistant, or both, i.e., WGAR/OuaR) of MDAY-D2, were examined for the presence of a tumor-associated antigen (TAA). Of 15 mice examined, tumor antigen-loss variants were detected in only 1 animal. These antigen-loss metastases arose in a mouse injected with the WGAR variant called MDW4. The tumor at the site of inoculation retained the TAA, whereas all four of the metastases removed from liver, spleen and other tissues were antigen-loss variants. The antigen-loss variants were not killed by cytotoxic T cells (CTL) directed against the TAA of the parental tumor, did not competitively inhibit CTL lysis of MDW4 targets in a 'cold target' inhibition test, and were not able to elicit a CTL response. In vivo immunization-protection (challenge) experiments also showed that the metastases did not express the TAA of MDAY-D2. Unlike the WGAR phenotypes, which were lost in all spontaneous metastases recovered from MDW4-injected mice, loss of the TAA appeared to be an uncommon event. Antigen-loss tumor cell variants are discussed in terms of their relevance to metastasis, and in regard to their use in the study of T cell-mediated cytotoxicity of tumor cell populations.

Animals

Molecular genetics of glucose-6-phosphate dehydrogenase (G6PD) deficiency in Spain: identification of two new point mutations in the G6PD gene.

In order to explore the nature of glucose-6-phosphate dehydrogenase (G6PD) deficiency in Spain, we have analysed the G6PD gene in 11 unrelated Spanish G6PD-deficient males and their relatives by using the polymerase chain reaction and single-strand conformation polymorphism (PCR-SSCP) analysis combined with a direct PCR-sequencing procedure and PCR-restriction enzyme (RE) analysis. We have identified eight different missense mutations, six of which have been reported in previously described G6PD variants. In nine patients who had presented with acute favism we found the following mutations: G6PD A-376G-202A (four cases), G6PD Union1360T (two cases), G6PD Mediterranean563T (one case) and G6PD Aures143C (one case). In the remaining patient a novel A to G transition was found at nucleotide position 209 which has not been reported in any other ethnic group. This mutation results in a (70) Tyr to Cys substitution and the resulting G6PD variant was biochemically characterized and designated as G6PD Murcia. This new mutation creates a Bsp 1286I recognition site which enabled us to rapidly detect it by PCR-RE analysis. In two patients with chronic non-spherocytic haemolytic anaemia (CNSHA) we found the underlying genetic defects, as had been noted previously, to be located within a cluster of mutations in exon 10. One of them had the T to C transition at nucleotide 1153, causing a (385) Cys to Arg substitution, previously described in G6PD Tomah. The other, previously reported as having a variant called G6PD Clinic, has a G to A transition at nucleotide 1215 that produces a (405) Met to Ile substitution, thus confirming that G6PD Clinic is a new class I variant.

Acute Disease

Molecular and cell isoforms during development.

Development proceeds by way of a discrete yet overlapping series of biosynthetic and restructuring events that result in the continued molding of tissues and organs into highly restricted and specialized states required for adult function. Individual molecules and cells are replaced by molecular and cellular variants, called isoforms; these arise and function during embryonic development or later life. Isoforms, whether molecular or cellular, have been identified by their structural differences, which allow separation and characterization of each variant. These isoforms play a central and controlling role in the continued and dynamic remodeling that takes place during development. Descriptions of the individual phases of the orderly replacement of one isoform for another provides an experimental context in which the process of development can be better understood.

Actins

Molecular characterization of variant translocations in chronic myelogenous leukemia.

Five to ten percent of the Ph-positive cases of chronic myelogenous leukemia (CML), termed variant translocations, involve at least one chromosome in addition to 9 and 22 in the abnormality. The involvement of chromosome 9 band q34, where the c-abl oncogene has been localized, is not always cytogenetically detectable in so called variant translocations due to complex rearrangements. We present two cases having the most frequently involved chromosomes (#3 and #17) in such translocations. In one case, both chromosome 9's were cytogenetically normal while in the other, band 9q34 was so called 'masked' or 'hidden'. After molecular evaluation using in situ hybridization and Southern blotting techniques, the involvement of the altered bcr/abl gene was demonstrated and the cytogenetic analysis was revised. Utilization of molecular probes in the evaluation of such cases should become a routine diagnostic procedure in detecting the exchange of bcr and c-abl sequences.

Adult

Genetic variants of alpha-1-antitrypsin (AAT).

This paper reviews the genetic variants of alpha-1-antitrypsin (AAT) which have been sequenced with special emphasis on the s.c. deficiency variants. These result in AAT low plasma levels via three main mechanisms: 1) intracellular storage; 2) intracellular degradation; 3) lack of synthesis. Intracellular storage occurs with the classical Z variant and with a few variants called M-like, because of their isoelectric focusing (IF) pattern. The storage phenomenon causes liver damage and can be demonstrated at both light and electron microscopic level with the help of immunohistochemistry. We report a new deficiency variant of AAT (M-Cagliari) characterized by very low plasma levels, massive storage of AAT and liver cirrhosis. By using immunohistochemical techniques and DNA analysis we could demonstrate that M-Cagliari has antigenic and genetic properties other than the Z AAT.

Base Sequence

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

A de novo algorithm for allele reconstruction from Oxford nanopore amplicon reads, with application to CYP2D6.

MOTIVATION: The Oxford Nanopore Technologies' sequencing platform offers a path towards bedside genomics, producing long reads that can completely cover a gene of interest, and detect any known or novel variant the gene contains. However, the analysis of these long reads to identify actionable genotypes remains challenging and typically requires customization depending on the target gene. RESULTS: Here, we describe a generic algorithm to accurately reconstruct allele sequences derived from long-reads of amplicon-based data. Rather than calling variants directly from these long-reads, our method takes a "sequence-first" approach, performing an unbiased reconstruction of the underlying amplicon sequences to generate high-confidence reconstructed allele sequences. This is done without user input of the target gene, allowing for any source amplicon to be reconstructed. These high-confidence reconstructed allele sequences are then compared to the genomic reference sequence of the gene to infer the specific diplotype present in the sample. This approach is agnostic towards the number of genes and alleles present and readily detects novel variants. We demonstrate our approach using three independent data sets for CYP2D6, a diverse and complex gene with over 175 known alleles of clinical significance. We show how our approach can accurately recover validated CYP2D6 diplotypes from 20 Coriell samples covering 14 distinct alleles, using different amplicons, flow cell versions, and depths. This includes inferring occurrences of allele duplication events from relative abundances of each allele, a critical factor for ascribing functional effects to a diplotype. Further, we demonstrate our approach's utility for other genomic regions, including HLA. AVAILABILITY: Custom code is available at the following GitHub repository, along with instructions for use and test data: https://github.com/scottdbrown/allele-reconstruction-long-read-amplicon-data. A snapshot of the code at the time of publication is available on Zenodo.org; doi 10.5281/zenodo.19716004. Raw .fastq sequence data for our three sequencing runs is available at the SRA under Bioproject PRJNA1357883 (https://www.ncbi.nlm.nih.gov/bioproject/1357883).

Alleles

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article

The Philadelphia chromosome. Considerations based on studies of variant Ph translocations.

The nature of the Philadelphia (Ph) translocation and the process of its formation were studied by attempting various chromosome banding analyses of variant Ph translocations among 210 patients with Ph-positive chronic myelocytic leukemia examined at the National Institute of Radiological Sciences, Chiba. The following assumptions could be drawn from the results of the analyses: 1) The involvement of specific regions of chromosomes #9 and #22, q34 and q11, respectively, is an indispensable condition of the Ph translocation. 2) The so-called variant Ph translocations are all complex and are derived from a standard Ph translocation. 3) The Ph translocations, both standard and complex ones, are not always stable. The complex translocations are subject to further chromosome evolution, as is the conversion of the standard translocation to complex translocations. There seems to be no fundamental difference between the standard and complex Ph translocations, with the latter being merely a more progressed form of the former. Analyses at the molecular level of the same cases employed in this study are yielding results that support the above assumptions.

Chromosome Banding

Interferon-alpha: a gene family in therapeutic use.

Several variants of interferon-alpha (IFN-alpha) were isolated and purified to homogeneity. They differed to various degrees in biological properties. However, three IFN-alpha 2 variants showed only minor differences from a variant called IFN-alpha 88 with regard to their ability to inhibit growth and to bind to specific receptors, tested on Daudi cells. Two monoclonal antibodies were studied, which showed overlapping specificity for at least one peptide obtained after HPLC separation of tryptic digests. The monoclonal antibodies could discriminate between sequence differences to a much higher degree than the receptor on Daudi cells. It is concluded that the receptor is degenerate and binds well to different structural variants of IFN and that for therapeutic use, several of the variants will probably have the same biological potency.

Antibodies, Monoclonal

Vcfexpress: flexible, rapid user-expressions to filter and format VCFs.

MOTIVATION: Variant call format (VCF) files are the standard output format for various software tools that identify genetic variation from DNA sequencing experiments. Downstream analyses require the ability to query, filter, and modify them simply and efficiently. Several tools are available to perform these operations from the command line, including BCFTools, vembrane, slivar, and others. RESULTS: Here, we introduce vcfexpress, a new, high-performance toolset for the analysis of VCF files, written in the Rust programming language. It is nearly as fast as BCFTools, but adds functionality to execute user expressions in the lua programming language for precise filtering and reporting of variants from a VCF or BCF file. We demonstrate performance and flexibility by comparing vcfexpress to other tools using the vembrane benchmark. AVAILABILITY AND IMPLEMENTATION: vcfexpress is available under the MIT license at https://github.com/brentp/vcfexpress with code used for the manuscript deposited in https://doi.org/10.5281/zenodo.14756838.

Software