Search PubMedSearch

SEARCH · Search PubMed

Results for “Genetic testing algorithm”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Large-scale proteomics profiling of peripheral blood of DM1 patients identifies biomarkers for disease severity and functional capacity.

BackgroundMyotonic Dystrophy Type 1 (DM1), the most common genetic neuromuscular disorder in adults, poses significant challenges for drug development due to its multisystem nature and high clinical variability in symptoms and disease progression. With a growing number of therapies entering clinical trials, this study addresses the urgent need for biomarkers that can serve as surrogate endpoints.MethodsWe profiled 437 serum samples from adult DM1 patients collected at two timepoints of the OPTIMISTIC trial using bottom-up mass spectrometry with data-independent acquisition. Associations between protein expression, the disease-causing CTG-repeat and 25 clinical outcome measures were studied using linear mixed-effect models. All key study findings were validated in an independent cohort of 69 DM1 patients and 10 healthy controls.ResultsOf the 259 identified proteins, 161 showed significant associations with the CTG-repeat length (FDR&#x2009;<&#x2009;5%). Hypogammaglobulinemia was confirmed and shown to be worse in severely affected patients. A strong proteomic signature was associated with clinical measures of functional capacity, with the 6-Minute Walk Test showing the strongest signal (70 associations, FDR&#x2009;<&#x2009;5%). These novel associations reveal a compelling link between chronic inflammation and reduced functional capacity. A machine learning algorithm identified a minimal set of 13 proteins robustly reflecting both the underlying genetic defect and functional capacity.ConclusionsDM1 induces a broad disease fingerprint in the serum proteome, predominantly affecting proteins of the immune system. A carefully selected panel of proteins showed the greatest potential to meet the statistical criteria required for surrogate endpoints in clinical trials.

Humans

Segregation analysis of quantitative traits in nuclear families: comparison of three program packages.

Segregation analysis frequently is used to test for the presence of major gene effects and to estimate the various genetic and environmental components contributing to diseases. Recent advances in both theoretical models and computational algorithms have provided a number of new programs for performing segregation analyses. We compared two newer programs: REGC (part of the package "SAGE") and FISHER/MENDEL with an older established program (PAP) to determine relative accuracy in recovering parameter values and asymptotic standard errors, ability to discriminate between alternative transmission models, and execution speeds. Each program was applied to a set of computer simulations of a quantitative trait generated under a variety of genetic models. The results of these comparisons indicated that all the programs provided very similar parameter estimates, but that they differed in their abilities to identify the correct mode of transmission. In our simulations, PAP more often led to the selection of the correct transmission model, whereas REGC frequently indicated the presence of a major gene in simulations of purely polygenic transmission. Relative speeds for the programs differed, and their rank ordering varied with the complexity of the model being fitted. Although REGC was the fastest program for fitting a major gene or mixed model, it was by far the slowest program for estimating parameters in a sporadic or polygenic model.

Computer Simulation

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to detect CNVs, which can be used in genetic association tests. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (those present on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has led to excessive reduction in overall sensitivity since arrays can have an undesirably low probe overlap. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4906 individuals genotyped across three different arrays, we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g. use of consensus probes only). We further demonstrate that MarkerMatch matches the CNV detection from current practice in terms of F1 score and PPV for larger CNVs. We also optimize MarkerMatch parameters, DMAX and Method, and find an optimal DMAX setting at 10&#x2009;kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis. AVAILABILITY: The R package for MarkerMatch is available at: https://github.com/FranjoIM/MarkerMatch. The code used for analysis and implementation is available at: https://doi.org/10.5281/zenodo.18460979. The live notebook is available at https://fivankovic.notion.site/2026-markermatch.

DNA Copy Number Variations

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

A strategy for using multiple linked markers for genetic counseling.

A strategy for using multiple linked markers for genetic counseling is to test sequentially individual markers until a diagnosis can be made. We show that in order to minimize the number of tests performed per case while diagnosing all informative cases the order in which the markers are to be tested is critical. We describe an algorithm to obtain this order using the parameter "I," the frequency of informative cases. The I value for a specific locus used depends on the marker frequency, association with the disease locus, and also on the informativeness of the marker loci already tested. Realizing that a direct assay for the beta S gene already exists, and that most cases of beta-thalassemia in Mediterraneans can be directly diagnosed using synthetic oligonucleotide probes, we illustrate the above technique by examining nine DNA polymorphisms in the human beta-globin cluster for their ability to diagnose sickle-cell anemia in American blacks and beta-thalassemia in Mediterraneans. This analysis shows that 95.39% of all sickle-cell pregnancies can be diagnosed by testing a subset of only six markers chosen by our algorithm. Furthermore, six markers can also diagnose 88.03% of beta-thalassemia in Greeks and 83.56% of beta-thalassemia in Italians. The test set is different from that suggested by the individual informative frequencies due to nonrandom associations between the restriction sites.

Anemia, Sickle Cell

Performing the exact test of Hardy-Weinberg proportion for multiple alleles.

The Hardy-Weinberg law plays an important role in the field of population genetics and often serves as a basis for genetic inference. Because of its importance, much attention has been devoted to tests of Hardy-Weinberg proportions (HWP) over the decades. It has long been recognized that large-sample goodness-of-fit tests can sometimes lead to spurious results when the sample size and/or some genotypic frequencies are small. Although a complete enumeration algorithm for the exact test has been proposed, it is not of practical use for loci with more than a few alleles due to the amount of computation required. We propose two algorithms to estimate the significance level for a test of HWP. The algorithms are easily applicable to loci with multiple alleles. Both are remarkably simple and computationally fast. Relative efficiency and merits of the two algorithms are compared. Guidelines regarding their usage are given. Numerical examples are given to illustrate the practicality of the algorithms.

Alleles

Trade-off between gestational age and miscarriage risk of prenatal testing: does it vary according to genetic risk?

It has been generally assumed that as genetic risk rises, so the higher procedure-related miscarriage rates of diagnostic tests done earlier in gestation become more acceptable. To test the hypothesis a decision tree was used, in which the only differences between two tests A and B were that A was carried out earlier in pregnancy and was more likely to cause miscarriage. Over a wide range of rankings for the three outcomes (procedure-related miscarriage of a normal baby, early termination of pregnancy after test A, late termination of pregnancy after test B), the expected utility (relative desirability) of an earlier, but more risky, test was greater at a high (1 in 4) than at a low (1 in 100) genetic risk.

Abortion, Spontaneous

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R&#xb2;=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans

MarkerMatch: A Proximity-Based Probe-Matching Algorithm for Joint Analysis of Copy-Number Variants from Different Genotyping Arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to call CNVs, which can be used in association tests, such as association between CNV number and disease status. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (the intersection encompassing the probes that occur in common on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has, however, led to excessive reduction in overall sensitivity of CNV calls as arrays can have an undesirably low overlap of probe sets. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4,906 individuals genotyped across three different arrays (Global Screening Array, Omni2.5 array, and Omni Express Exome array), we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g., use of consensus probes only). We further demonstrate that MarkerMatch exceeds the output from current practice in terms of F1 score, Fowlkes-Mallows index, and Jaccard index. We also optimize MarkerMatch parameters, D MAX and Method, and find an optimal D MAX setting at 10kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis.

Journal Article

Development and validation of blood-based diagnostic biomarkers for Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) using EpiSwitch&#xae; 3-dimensional genomic regulatory immuno-genetic profiling.

Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS) is a debilitating, multifactorial disorder characterised by profound fatigue, post-exertional malaise, cognitive impairments, and autonomic dysfunction. Despite its significant impact on quality of life, ME/CFS lacks definitive diagnostic biomarkers, complicating diagnosis and management. Recent evidence highlights potential blood tests for ME/CFS biomarkers in immunological, genetic, metabolic, and bioenergetic domains. Chromosome conformations (CCs) are potent epigenetic regulators of gene expression and cross-tissue exosome signalling. We have previously developed an epigenetic assay, EpiSwitch&#xae;, that employs an algorithm-based CCs analysis. Using EpiSwitch&#xae; technology, we have shown the presence of disease-specific CCs in peripheral blood mononuclear cells (PBMCs) of patients with amyotrophic lateral sclerosis (ALS), rheumatoid arthritis (RA), prostate and colorectal cancers, diffuse Large B-cell lymphoma and severe COVID-19. In a recent paper, we have identified a profile of systemic chromosome conformations in cancer patients reflective of the predisposition to respond to immune checkpoint inhibitors, PD-1/PD-L1 antagonists, with 85% accuracy. In this Retrospective case/control study (EPI-ME, Epigenetic Profiling Investigation in Myalgic Encephalomyelitis), we used whole blood samples retrospectively collected from n&#x2009;=&#x2009;47 patients with severe ME/CFS and n&#x2009;=&#x2009;61 age-matched healthy control patients to perform whole-genome 3D DNA screening for CCs correlating to ME/CFS diagnosis. We identified a 200-marker model for ME/CFS diagnosis (Episwitch&#xae;CFS test). First testing on the retrospective independent validation cohort demonstrated a strong systemic ME/CFS signal with a sensitivity of 92% and a specificity of 98%.Pathways analysis revealed several likely contributors to the pathology of ME/CFS, including interleukins, TNF&#x3b1;, neuroinflammatory pathways, toll-like receptor signalling and JAK/STAT. Comparison with pathways involved in the action of Rituximab and glatiramer acetate (Copaxone) (therapies with potential in ME/CFS treatment) identified IL2 as a shared pathway with clear patient clustering, indicating a possibility of a potential responder group for targeted treatment.

Humans

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

The Elston-Stewart algorithm for continuous genotypes and environmental factors.

The Elston-Stewart algorithm for a normally distributed trait under a polygenic model is explained in detail and extended to allow for other continuous environmental variables. This formulation is especially useful for large pedigrees, as it avoids the need to invert matrices. Whereas it may not be feasible by this method to estimate all the various components of previously suggested models for polygenic inheritance, it can allow for a reasonably flexible pedigree correlational structure under which valid tests can be performed for fixed effects that may affect the phenotype.

Algorithms

Pharmacokinetic-pharmacogenetic modelling in the detection of polymorphisms in xenobiotic metabolism.

Study of the genetic control of xenobiotic metabolism is hindered in the areas of detecting new polymorphisms and estimating the frequency and enzyme activity of each phenotype. Using computer simulation we have looked at pharmacogenetic-pharmacokinetic models based on two alleles under Hardy-Weinberg equilibrium. The distributions of the area under the concentration-time curve (AUC) or urinary ratios were modelled and the effects of incomplete urine collections, sequential, parallel and non-linear pathways investigated. The statistical methods for the detection of bimodality in these distributions were explored. The drug/metabolite ratio, which has a good theoretical basis, was confirmed to be most sensitive and robust to changes in bioavailability, urinary excretion, hepatic blood flow and variation in non-polymorphic enzyme activity but not parallel, sequential or non-linear routes of metabolism. Graphical methods, while able to illustrate deviations from normality, were not specific in detecting bimodality and the hypothesis testing methods were found to be heavily dependent upon their assumptions.

Algorithms

Evidence of hyperplanes in the genetic learning of neural networks.

Genetic Algorithms have been successfully applied to the learning process of neural networks simulating artificial life. In previous research we compared mutation and crossover as genetic operators on neural networks directly encoded as real vectors (Manczer and Parisi 1990). With reference to crossover we were actually testing the building blocks hypothesis, as the effectiveness of recombination relies on the validity of such hypothesis. Even with the real genotype used, it was found that the average fitness of the population of neural networks is optimized much more quickly by crossover than it is by mutation. This indicated that the intrinsic parallelism of crossover is not reduced by the high cardinality, as seems reasonable and has indeed been suggested in GA theory (Antonisse 1989). In this paper we first summarize such findings and then propose an interpretation in terms of the spatial correlation of the fitness function with respect to the metric defined by the average steps of the genetic operators. Some numerical evidence of such interpretation is given, showing that the fitness surface appears smoother to crossover than it does to mutation. This confirms indirectly that crossover moves along privileged directions, and at the same time provides a geometric rationale for hyperplanes.

Algorithms

Multifactorial genetic models for quantitative traits in humans.

Quantitative traits measured in human families can be analyzed to partition the total population variance into genetic and environmental components, or to elucidate the genetic mechanism involved. We review the estimation of variance components directly from human pedigree data, or in the form of path coefficients from correlations between pairs of relatives. To elucidate genetic mechanisms, a mixed model that allows for segregation at a major locus, a polygenic effect and a sibling environmental correlation is described for nuclear families. In each case appropriate likelihoods are derived as a basis, using numerical maximum likelihood methods, for parameter estimation and hypothesis testing. A general model is then described that allows for several familial sources of environmental variation, assortative mating, and both major gene and polygenic effects; and an algorithm for calculating the likelihood of a pedigree under this model is indicated. Finally, some of the remaining problems in this area of biometric analysis are pointed out.

Genetic Variation

Adapting best linear unbiased prediction (BLUP) for timely genetic evaluation: II. Progeny traits in multiple contemporary groups within a herd.

The second step of a procedure to partially circumvent the voluminous calculations for some BLUP (Best Linear Unbiased Prediction) computing algorithms for genetic evaluation is presented. In addition, the procedure allows timely evaluations of each contemporary group. This procedure is pertinent especially for polytocous species such as swine and poultry, for which the occurrence of full-sib families makes the inclusion of dam effects in the model necessary and tests are completed throughout the year. Formulas are developed for a model including sires, dams, individuals within full-sib families and records within individuals. This model has a fundamentally hierarchical structure but includes some cross-classification. The formulas for predictors combine information across contemporary groups within a herd and incorporate relationships between sires and(or) dams in that herd. Formulas to approximate prediction error variances also are developed.

Algorithms

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding