Search PubMedSearch

SEARCH · Search PubMed

Results for “high-throughput genotyping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genetic analysis of partial duplication of the long arm of chromosome 16.

BACKGROUND: Pure partial trisomy 16q12.1q22.1 is a rare chromosome copy number variant (CNV). The primary clinical phenotypes associated with this syndrome include abnormal facial morphology, global developmental delay (GDD), short stature, and reported predisposing factors for atypical behavior, autism, the development of learning disabilities, and neuropsychiatric disorders. The dosage-sensitive genes associated with partial trisomy are not disclosed preventing to establish a genotype-phenotype correlation. METHODS: We report a case of a Chinese patient diagnosed with GDD and an abnormal facial shape, who was found to have partial trisomy 16 through karyotyping and high-throughput sequencing analysis. Karyotype and CNV tracing analyses were also conducted on the biological parents of the patient to assess for any chromosomal structural abnormalities. Additionally, we included 29 patients with pure partial trisomy 16q, reported in the DECIPHER database and the literature. We and performed a genotype-phenotype correlation analysis. RESULTS: The proband, a 2-year-old female, was found to have a de novo 21.96 Mb duplication located between 16q12.1q22.1, with no other deletions observed on other chromosomes, indicating a pure partial trisomy of 16q. Through genotype and phenotype analysis of 29 individuals, we found that patients with the duplicated region located at the distal region of 16q may exhibit more severe symptoms than those with duplication at the proximal region; however, no relationship was identified between phenotype and the size of the duplicated segment. CONCLUSION: We report, for the first time, a patient with partial trisomy 16q validated by multiple genetic tests, including CNV-seq, whole exome sequencing (WES), and karyotyping. It is speculated that partial trisomy of 16q may be associated with continuous gene duplication. However, functional studies are necessary to identify the causative gene or critical region linked to duplication syndrome of chromosome 16q.

Child, Preschool

Application of third-generation sequencing technology for identifying rare α- and β-globin gene variants in a Southeast Chinese region.

BACKGROUND: Third-generation sequencing (TGS) based on long-read technology has been gradually used in identifying thalassemia and hemoglobin (Hb) variants. The aim of the present study was to explore genotype varieties of thalassemia and Hb variants in Quanzhou region of Southeast China by TGS. METHODS: Included in this study were 6,174 subjects with thalassemia traits from Quanzhou region of Southeast China. All of them underwent common thalassemia gene testing using the DNA reverse dot-blot hybridization technology. Subjects who were suspected as rare thalassemia carriers were further subjected to TGS to identify rare or novel α- and β-globin gene variants, and the results were verified by Sanger sequencing and/or gap PCR. RESULTS: Of the 6,174 included subjects, 2,390 (38.71%) were identified as α- and β-globin gene mutation carriers, including 40 carrying rare or novel α- and β-thalassemia mutations. The αCD30(-GAG)α and Hb Lepore-Boston-Washington were first reported in Fujian province Southeast China. Moreover, the βCD15(TGG> TAG), βIVS-II-761, β0-Filipino(~ 45 kb deletion), and Hb Lepore-Quanzhou were first identified in the Chinese population. In addition, 35 cases of Hb variants were detected, the rare Hb variants of Hb Jilin and Hb Beijing were first reported in Fujian province of China. Among them, one case with compound αααanti3.7 and Hb G-Honolulu variants was identified in this study. CONCLUSION: Our findings may provide valuable data for enriching the spectrum of thalassemia and highlight the clinical application value of TGS-based α- and β-globin genetic testing.

Humans

Evaluation of bone preparation approaches using length-based analysis and targeted sequencing for forensic human identification of historic skeletal remains.

Advances in DNA technology have significantly enhanced the forensic community's ability to develop genetic profiles from unidentified human skeletal remains. However, sampling requires mechanical grinding of hard tissues before DNA isolation. This processing can compromise genetic profiles, particularly in aged bones. We compared the industry-standard pulverization method with an alternative powder-free preparation involving prolonged demineralization and subsequent slicing of 19th-century cortical bone. Data from DNA quantification, STR genotyping, and targeted SNP sequencing were used to evaluate powdered samples versus demineralized slices from paired human bones. Average human DNA yields for pulverized samples and demineralized slices were 0.032&#x2009;ng and 0.692&#x2009;ng, respectively. Demineralized slices recovered more amplifiable DNA than traditional homogenization methods (p&#x2009;<&#x2009;0.05). No pulverized samples produced STR profiles, whereas demineralized slices from the same bone samples yielded partial profiles. Samples underwent DNA repair, library preparation, and hybridization capture using the FORensic Capture Enrichment (FORCE) panel. Applying low-coverage (1X) analysis of high-throughput sequencing (HTS) data, demineralized slices outperformed those prepared by traditional pulverization methods (p&#x2009;<&#x2009;0.05) and substantially increased the information recovered compared with conventional STR analysis methods. Based on HTS data from pulverized samples, DNA fragment length ranged from 27 to 95&#x2009;bp, and FORCE SNP recovery was 33.23%. In contrast, for demineralized slices, DNA fragment length ranged from 85 to 114&#x2009;bp, and FORCE SNP recovery was 83.24%. The required reagents and equipment are typically available in forensic labs, and the workflow outlined herein significantly increases the success of DNA recovery from challenging skeletal samples.

Humans

Analysis of HLA Allelic and Haplotypic Frequencies in a Cohort of Bone Marrow Donors in Catalonia: Impact of Next-Generation Sequencing on Donor Registry Quality.

In haematopoietic stem cell transplantation (HSCT), the volunteer unrelated donor (VUD) has become the most common strategy in Europe, as improved outcomes are achieved through HLA compatibility at allelic-level resolution. In this context, the implementation of next-generation sequencing (NGS) in histocompatibility typing laboratories has significantly enhanced the quality of bone marrow registries, enabling a high level of resolution at a lower cost and improved performance. In this study, we analyse a large cohort of 21,787 bone marrow donors in Catalonia and present the observed HLA allelic and haplotypic frequencies, along with their linkage disequilibria. HLA-A, -B, -C, -E and -G were genotyped at full resolution, while -DRB1, -DQB1, -DQA1, -DPA1 and -DPB1 were genotyped at high resolution. We identified 236 new officially named HLA alleles, both coding and non-coding regions. This study highlights that the implementation of high-throughput HLA typing has led to an increase in the number of registered donors and an improvement in quality, which has been reflected in a rise in the number of effective donors.

Humans

Order among chaos: High throughput MYCroplanters can distinguish interacting drivers of host infection in a highly stochastic system.

The likelihood that a host will be susceptible to infection is influenced by the interaction of diverse biotic and abiotic factors. As a result, substantial experimental replication and scalability are required to identify the contributions of and interactions between the host, the environment, and biotic factors such as the microbiome. For example, pathogen infection success is known to vary by host genotype, bacterial strain identity and dose, and pathogen dose. Elucidating the interactions between these factors in vivo has been challenging because testing combinations of these variables quickly becomes experimentally intractable. Here, we describe a novel high throughput plant growth system (MYCroplanters) to test how multiple host, non-pathogenic bacteria, and pathogen variables predict host health. Using an Arabidopsis-Pseudomonas host-microbe model, we found that host genotype and bacterial strain order of arrival predict host susceptibility to infection, but pathogen and non-pathogenic bacterial dose can overwhelm these effects. Host susceptibility to infection is therefore driven by complex interactions between multiple factors that can both mask and compensate for each other. However, regardless of host or inoculation conditions, the ratio of pathogen to non-pathogen emerged as a consistent correlate of disease. Our results demonstrate that high-throughput tools like MYCroplanters can isolate interacting drivers of host susceptibility to disease. Increasing the scale at which we can screen drivers of disease, such as microbiome community structure, will facilitate both disease predictions and treatments for medicine and agricultural applications.

Arabidopsis

Deletion of the HLA-B Gene in One of the Inherited Haplotypes in a Northern&#x2009;European Family.

Targeted next generation sequencing-based HLA typing of a 17-year-old female transplant patient showed homozygosity for the HLA-B allele. The segregation analysis of HLA haplotypes of family members only allowed the conclusion that the B-allele was deleted in the haplotype inherited from the father and accordingly paternal grandfather, resulting in false homozygous genotyping. The subsequent whole-genome sequencing of the patient and her father confirmed an approximately 85&#x2009;kb deletion at 6p21.33 from the 5' end of the HLA-B to the 3' end of the HLA-C gene extending telomeric to HLA-C.

Adolescent

Missense variants pathogenicity annotation from homologous proteins.

MOTIVATION: High-throughput DNA sequencing has revealed millions of single nucleotide variants (SNVs) in the human genome, with a small fraction linked to disease. The effect of missense variants, which alter the protein sequence, is particularly challenging to interpret due to the scarcity of clinical annotations and experimental information. While using conservation and structural information, current prediction tools still struggle to predict variant pathogenicity. In this study, we explored the pathogenicity of homologous missense variants-variants in equivalent positions across homologous proteins-focusing on proteins involved in autosomal dominant diseases. RESULTS: Our analysis of 2976 pathogenic and 17&#xa0;555 non-pathogenic homologous variants demonstrated that pathogenicity can be extrapolated with 95% accuracy within a family, or up to 98% for closer homologs. Remarkably, the evaluation of 27 commonly used mutation predictor methods revealed that they were not fully capturing this biological feature. To facilitate the exploration of homologous variants, we created HomolVar, a web server that computationally predicts the pathogenesis of missense variants using annotations from homologous variants, freely available at https://rarevariants.org/HomolVar. Overall, these findings and the accompanying tool offer a robust method for predicting the pathogenicity of unannotated variants, enhancing genotype-phenotype correlations, and contributing to diagnosing rare genetic disorders. AVAILABILITY AND IMPLEMENTATION: HomolVar is freely available at https://rarevariants.org/HomolVar.

Mutation, Missense

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8&#xa0;&#xb0;C-1.4&#xa0;&#xb0;C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN

Identification of intragenic variants in pediatric patients with intellectual disability in Peru.

BACKGROUND: Intellectual disability in Latin America can reach a frequency of 12% of the population, these may include nutritional deficiencies, exposure to toxic or infectious agents, and the lack of universal neonatal screening programs. In 90% of patients with intellectual disability, the etiology can be attributed to variants in the genome. OBJECTIVE: to determine intragenic variants in patients with intellectual disability between 5 and 18 years old at Instituto Nacional de Salud del Ni&#xf1;o. METHODS: It is a descriptive cross-sectional study with convenience sampling. A total of 124 children diagnosed with intellectual disability were selected based on psychological test results and availability for whole exome sequencing. In addition, a chromosomal analysis of 6.55&#xa0;M was performed on ten patients with a negative result in sequencing. Relative and absolute frequencies and measures of central tendency and dispersion were determined according to their nature. In addition, multiple linear regression and Poisson regression were used to determine the association between some clinical characteristics and the probability of occurrence in patients with positive results. RESULTS: The median age of the patients was 6.3 (IQR&#x2009;=&#x2009;5.95), males accounted for 57.3%, and 91.9% of the cases had mild intellectual disability. Exome sequencing determined the etiology in 30.6% of patients with intellectual disability, of which 52.6% were autosomal dominant inheritance. The most frequent genes found were MECP2, STXBP1 and LAMA2. A broad genotype-phenotype correlation was identified, highlighting the genetic heterogeneity of intellectual disability in this population. The presence of dermatologic lesions, dystonia, peripheral neurological disorders, and fourth finger flexion limitation were observed more frequently in patients with intellectual disability with "positive results". CONCLUSIONS: This study shows that one-third of patients with intellectual disability exhibit intragenic variants, highlighting the importance of genetic analysis for accurate diagnosis. The identification of genes such as MECP2, STXBP1, and LAMA2 underscores the genetic heterogeneity of intellectual disability in the studied population. These findings emphasize the need for genetic testing in clinical management and the implementation of early detection programs in Peru.

Humans

UPDhmm: detecting uniparental disomy from NGS trio data.

SUMMARY: Uniparental disomies (UPDs) are copy-neutral chromosomal alterations that occur when both copies of a chromosome pair (entire or segmental) come from one parent. UPDs, including isodisomies (identical parental chromosome) and heterodisomies (two different homologs from the same parent), reflect meiotic and/or mitotic aberrations of chromosomal segregation that can be associated with congenital or acquired disease. Despite their relevance, current methods to detect UPDs using sequence data (exomes or genomes) have limited sensitivity for small events, cannot precisely determine the UPD sub-type or coordinates, and perform poorly when including individuals or populations with consanguinity. We present UPDhmm, a novel tool that uses trio-based sequence data (proband and parents) and models inheritance patterns. UPDhmm predicts the most likely inheritance scenario, normal Mendelian inheritance versus UPD event, based on genotype combinations using a Hidden Markov Model (HMM). We validated the method using simulations on exome and genome data from 1000-Genomes projects. UPDhmm overperformed currently available methods in detecting simulated UPD events in both data types. We applied UPDhmm to a collection of nearly 2400 families with a proband with autism spectrum disorder (Simons Simplex Collection Project) and identified UPD events in two affected individuals, one of them previously unreported. These two events, a paternal isodisomy of chr8 and a maternal heterodisomy of chr22, can be genetic causes of the disease, demonstrating the clinical utility of UPDhmm. Thus, UPDhmm can facilitate the incorporation of UPD detection into clinical pipelines of genomic analysis. AVAILABILITY AND IMPLEMENTATION: UPDhmm is implemented in R and is available in the Bioconductor package (version 1.5.0): https://www.bioconductor.org/packages/release/bioc/html/UPDhmm.html. The source code can be found at https://github.com/martasevilla/UPDhmm under the MIT license.

Uniparental Disomy

Diagnostic and phylogenetic perspectives of the 2023 Murray Valley encephalitis virus outbreak in Australia: an observational study.

BACKGROUND: An outbreak of Murray Valley encephalitis virus (MVEV), the largest since 1974, was observed in Australia between Jan 1 and July 31, 2023. This study aims to characterise the utility of diagnostic platforms, testing algorithms, and genomic characteristics of MVEV to facilitate a comprehensive framework for MVEV testing and surveillance in the outbreak setting. METHODS: In this observational study, we assessed flavivirus diagnostics for all patients with suspected Murray Valley encephalitis in Australia from Jan 1 to July 31, 2023. We included all patients with confirmed Murray Valley encephalitis, probable Murray Valley encephalitis, or acute unspecified flavivirus infection using the Communicable Diseases Network Australia case definition. Cases were excluded if an alternative diagnosis was identified. We collected blood, serum, cerebrospinal fluid, brain tissue, urine, or a combination of these samples, as appropriate and at the discretion of the treating clinician. We conducted multimodal diagnostic testing, which included flavivirus-specific serological and nucleic acid amplification testing. Metagenomic next-generation sequencing, including next-generation deep sequencing, target-enrichment, and targeted amplification, was conducted on human and representative mosquito-derived samples obtained from established mosquito population surveillance programmes for phylogenetic analysis. FINDINGS: 27 patients with encephalitis were assessed for MVEV between Jan 1, 2023, and July 31, 2023, 23 (85%) of whom fulfilled national case definitions for confirmed Murray Valley encephalitis. Patient ages ranged from 6 weeks to 83 years (median 62&#xb7;0 years [IQR 31&#xb7;0-67&#xb7;5]) and patients were mostly male (21 [78%] male patients and six [22%] female patients). Incidence varied widely by geographical region and was highest in the Northern Territory (32&#xb7;0 per 1&#x2009;000&#x2009;000 population). Diagnostic specimen collection generally occurred promptly (median 6&#xb7;0 days [IQR 4&#xb7;0-14&#xb7;5] from symptom onset to diagnostic specimen collection). In seven patients, case assignation relied on convalescent serum samples to assess for seroconversion or an appropriate rise in antibody titre (to four times the initial value or greater), or both. MVEV-specific IgM was detectable in serum samples of 17 (81%) of 21 patients tested by day 7 and MVEV IgG or total antibody (TAb) were detected in 18 (100%) of 18 patients tested by day 30. MVEV-specific IgM (or TAb) and MVEV RNA were detected in cerebrospinal fluid collected within 14 days of symptom onset in nine (39%) of 23 patients and seven (28%) of 25 patients, respectively. Phylogenetic analysis revealed two circulating MVEV genotypes, G1A and G2, in mosquitoes and humans in 2023. In southeast Australia, only G1A was detected and probably introduced from enzootic foci in northern Australia. INTERPRETATION: This study provides a comprehensive overview of the diagnostic workflows and phylogenetic evaluations used during the 2023 MVEV outbreak in Australia, emphasising the importance of a multimodal approach for accurate and timely confirmation of flavivirus infection. Further One Health surveillance for MVEV and other zoonotic flaviviruses is key, given potential expanded ecological niches in the context of episodic climatic events. FUNDING: None.

Humans

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software

A de novo algorithm for allele reconstruction from Oxford nanopore amplicon reads, with application to CYP2D6.

MOTIVATION: The Oxford Nanopore Technologies' sequencing platform offers a path towards bedside genomics, producing long reads that can completely cover a gene of interest, and detect any known or novel variant the gene contains. However, the analysis of these long reads to identify actionable genotypes remains challenging and typically requires customization depending on the target gene. RESULTS: Here, we describe a generic algorithm to accurately reconstruct allele sequences derived from long-reads of amplicon-based data. Rather than calling variants directly from these long-reads, our method takes a "sequence-first" approach, performing an unbiased reconstruction of the underlying amplicon sequences to generate high-confidence reconstructed allele sequences. This is done without user input of the target gene, allowing for any source amplicon to be reconstructed. These high-confidence reconstructed allele sequences are then compared to the genomic reference sequence of the gene to infer the specific diplotype present in the sample. This approach is agnostic towards the number of genes and alleles present and readily detects novel variants. We demonstrate our approach using three independent data sets for CYP2D6, a diverse and complex gene with over 175 known alleles of clinical significance. We show how our approach can accurately recover validated CYP2D6 diplotypes from 20 Coriell samples covering 14 distinct alleles, using different amplicons, flow cell versions, and depths. This includes inferring occurrences of allele duplication events from relative abundances of each allele, a critical factor for ascribing functional effects to a diplotype. Further, we demonstrate our approach's utility for other genomic regions, including HLA. AVAILABILITY: Custom code is available at the following GitHub repository, along with instructions for use and test data: https://github.com/scottdbrown/allele-reconstruction-long-read-amplicon-data. A snapshot of the code at the time of publication is available on Zenodo.org; doi 10.5281/zenodo.19716004. Raw .fastq sequence data for our three sequencing runs is available at the SRA under Bioproject PRJNA1357883 (https://www.ncbi.nlm.nih.gov/bioproject/1357883).

Alleles