Search PubMedSearch

SEARCH · Search PubMed

Results for “Genetic testing algorithm”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Testing for bimodality in frequency distributions of data suggesting polymorphisms of drug metabolism--hypothesis testing.

1. The theory of methods of hypothesis testing in relation to the detection of bimodality in density distributions is discussed. 2. Practical problems arising from these methods are outlined. 3. The power of three methods of hypothesis testing was compared using simulated data from bimodal distributions with varying separation between components. None of the methods could determine bimodality until the separation between components was 2 standard deviation units and could only do so reliably (greater than 90%) when the separation was as great as 4-6 standard deviation units. 4. The robustness of a parametric and a non-parametric method of hypothesis testing was compared using simulated unimodal distributions known to deviate markedly from normality. Both methods had a high frequency of falsely indicating bimodality with distributions where the components had markedly differing variances. 5. A further test of robustness using power transformation of data from a normal distribution showed that the algorithms could accurately determine unimodality only when the skew of the distribution was in the range 0-1.45.

Computers

Diagnosis of twin zygosity by mailed questionnaire.

A deterministic questionnaire method for zygosity determination is developed for use in epidemiological studies of adult twins. It is based on the answers of both members of a twin pair to two questions on similarity and confusion in childhood. The algorithm of the method is used to determine the zygosity status of a twin pair at two different levels of certainty. The validity of the method is tested by making blood marker determinations of 11 polymorphic marker systems fro a random sample of 104 twin pairs. The agreement between questionnaire and blood marker diagnosis was 100%, but the stricter level of certainty left 8.7% in the nonclassified group. The genetical representativeness of the sample is tested by the allele distribution of the markers as compared to the Finnish population data as well as by the distribution of the number of intra-pair differences in blood markers.

Blood Group Antigens

A Monte Carlo method for combined segregation and linkage analysis.

We introduce a Monte Carlo approach to combined segregation and linkage analysis of a quantitative trait observed in an extended pedigree. In conjunction with the Monte Carlo method of likelihood-ratio evaluation proposed by Thompson and Guo, the method provides for estimation and hypothesis testing. The greatest attraction of this approach is its ability to handle complex genetic models and large pedigrees. Two examples illustrate the practicality of the method. One is of simulated data on a large pedigree; the other is a reanalysis of published data previously analyzed by other methods.

Algorithms

Nonrandom distribution of genotypes among red cell indices.

Venous blood samples were obtained from 25,302 healthy adults in Kentucky, USA. The red cell indices measured on these samples were evaluated by multiple stepwise regression analysis to derive an algorithm capable of discriminating the 138 individuals within this population who had genotypes AA, AC, AS or AA beta-thalassemia. The simple discriminant MCV2 x MCH with a cut-off set at 1530 detected 137 out of 138 of the heterozygotes with a false positive rate in this population of 4.4%. Other discriminants tested produced fewer false positives but also missed a sufficient number of heterozygotes to be unacceptable for genetic counselling purposes.

Discriminant Analysis

Integrating Next-Generation Sequencing into von Willebrand Disease Diagnostics: Insights from the PCM-EVW-ES Multicenter Project.

Von Willebrand disease (VWD) is the most common inherited bleeding disorder, caused by quantitative or qualitative defects in von Willebrand factor (VWF). Diagnosis is challenging and requires integrating bleeding history, VWF antigen and activity measurements, FVIII assays, and specialized phenotyping. Genetic testing is increasingly recognized as a key component. Here, we review current concepts in VWD diagnostics and highlight the Spanish Clinical and Molecular Profile of von Willebrand Disease (PCM-EVW-ES) project as a model for genomics-enabled precision medicine. PCM-EVW-ES is a multicenter initiative involving 48 hospitals, centralized phenotypic testing, and next-generation sequencing of the VWF coding region, enabling definitive classification in 730 individuals with VWD to date. Harmonized recruitment criteria and standardized workflows improve subtype assignment, uncover complex genotypes, refine genotype-phenotype correlations, and facilitate the identification of asymptomatic carriers. The PCM-EVW-ES variant spectrum highlights recurrent disease-causing variants in Spain and underscores the value of coordinated national registries for variant curation. Building on these data, we propose a diagnostic algorithm in which bleeding assessment and first-line VWF/FVIII assays, combined with, early VWF molecular testing increases diagnostic accuracy and guides targeted second-line investigations to confirm and refine VWD subtype classification. We also outline persisting challenges, including the interpretation of variants of uncertain significance and patients without identifiable pathogenic VWF variants, and future directions integrating third-generation sequencing, expanded gene panels, functional studies, and artificial-intelligence-driven multiomic approaches. Together, these advances illustrate how robust multicenter studies can bridge the gap between complex diagnostics and clinical practice in VWD.

Humans

DeepLabCut-based automated system reveals diverse temperature tolerance among medaka strains and related Oryzias species.

Temperature is a critical environmental factor influencing the physiology and behavior of ectothermic animals, yet conventional methods for evaluating thermal tolerance in fish rely on subjective manual observation of loss of equilibrium (LOE), limiting experimental throughput and introducing observer bias. Here, we developed an automated temperature tolerance evaluation system integrating DeepLabCut-based pose estimation with custom image processing algorithms to objectively quantify the timing of LOE during thermal stress tests. Our system incorporated region partitioning and color transformation preprocessing to improve keypoint detection accuracy, followed by a classification model combining ResNet34-based frame features with keypoint coordinates to objectively determine the timing of LOE without manual observation. Validation against manual annotation showed that the automated system achieved an accuracy comparable to the natural variability between trained investigators, and outperformed naive human observers, supporting its validity as an objective and reproducible alternative to manual scoring. Using this system, we characterized cold and heat tolerance across six medaka strains (Oryzias latipes: d-rR/TOKYO, HB11A, OK-Cab, HO5 and HdrR-II1; O. sakaizumii: HNI-II). Cold and heat tolerance assessment revealed inter-strain variation, with HdrR-II1 among the most cold- and heat-tolerant strains and HNI-II the least tolerant of both cold and heat stress. We further evaluated cold tolerance in medaka-related species (O. sinensis, O. cabaranensis, O. curvinotus, O. luzonensis, O. celebensis, and O. javanicus) and zebrafish (Danio rerio), revealing substantial interspecific variation that broadly corresponded with latitudinal distribution. O. latipes, distributed at the highest latitudes among the tested species, exhibited the greatest cold tolerance, whereas O. celebensis, O. javanicus, and other tropical or low-latitude species showed comparatively low cold tolerance. Our automated system provides a robust, high-throughput platform for thermal tolerance evaluation and, combined with the genetic and genomic resources available in medaka, establishes a foundation for elucidating the molecular mechanisms underlying temperature adaptation in fish.

Animals

A computer algorithm for testing potential prokaryotic terminators.

The nucleotide sequences of 30 factor-independent terminators of transcription with RNA polymerase from E. coli have been compiled and analyzed. The standard features - a stretch of thymine residues and a preceding dyad symmetry - are shared by most sequences, but there are striking exceptions which indicate that these features alone are not sufficient to describe these sites. In two thirds of the sequences the 3'-half of the dyad symmetry contains the pentanucleotide CGGG (G/C) or a close derivative; about one third have TCTG or a close derivative just downstream of the termination point. The TCTG -box might be implied in termination of stringently controlled operons of E. coli. An algorithm to locate terminators in templates of known nucleotide sequence has been constructed on the basis of correlation to the distribution of dinucleotides along the aligned signal sequences. The algorithm has been tested on natural sequences of a total length of about 11,500 N. It finds all known independent terminators and only a few other sites, including some of the rho-dependent and putative terminators.

Base Sequence

Estimating allele frequencies of hypervariable DNA systems.

Several polymorphisms of human DNA have been shown to be hypervariable due to the recurrence of a variable number of tandem repeats (VNTRs) in the lengths of allelic restriction fragments. The recurrence of allelic variants in this novel class of polymorphisms seems to comply well with a model of continuous random variables. Based on this assumption, we have compiled some simple algorithms for classification of continuous data and estimation of classes of relative frequencies and have implemented these routines for the management of databases storing hypervariable single locus DNA genetic systems. The algorithms are compiled in BASIC language and can be incorporated in task-oriented computer programs. Three procedures are discussed, based in turn on: (a) using predetermined, arbitrary classes; (b) point estimations of frequencies for single fragments using error measurements associated with the kilobase value assignment; (c) estimates of phenotype frequencies according to error measurements. Error measurements are obtained from a statistic of values pertaining to several restriction fragments (genomic controls) repeatedly tested in different experiments. Problems related to these approaches are discussed.

Algorithms

Phenotypic presentation of Mendelian disease across the diagnostic trajectory in electronic health records.

PURPOSE: To investigate the phenotypic presentation of Mendelian disease across the diagnostic trajectory in the electronic health record (EHR). METHODS: We applied a conceptual model to delineate the diagnostic trajectory of Mendelian disease to the EHRs of patients affected by 1 of 9 Mendelian diseases. We assessed data availability and phenotype ascertainment across the diagnostic trajectory using phenotype risk scores and validated our findings via chart review of patients with hereditary connective tissue disorders. RESULTS: We identified 896 individuals with genetically confirmed diagnoses, 216 (24%) of whom had fully ascertained diagnostic trajectories. Phenotype risk scores increased following clinical suspicion and diagnosis (P < 1&#xa0;&#xd7; 10-4, Wilcoxon rank sum test). We found that of all International Classification of Disease-based phenotypes in the EHR, 66% were recorded after clinical suspicion, and manual chart review yielded consistent results. CONCLUSION: Using a novel conceptual model to study the diagnostic trajectory of genetic disease in the EHR, we demonstrated that phenotype ascertainment is, in large part, driven by the clinical examinations and studies prompted by clinical suspicion of a genetic disease, a process we term diagnostic convergence. Algorithms designed to detect undiagnosed genetic disease should consider censoring EHR data at the first date of clinical suspicion to avoid data leakage.

Humans

gm: a practical tool for automating DNA sequence analysis.

The gm (gene modeler) program automates the identification of candidate genes in anonymous, genomic DNA sequence data. gm accepts sequence data, organism-specific consensus matrices and codon asymmetry tables, and a set of parameters as input; it returns a set of models describing the structures of candidate genes in the sequence and a corresponding set of predicted amino acid sequences as output, gm is implemented in C, and has been tested on Sun, VAX, Sequent, MIPS and Cray computers. It is capable of analyzing sequences of several kilobases containing multi-exon genes in less than 1 min execution time on a Sun 4/60.

Algorithms

The power of the N-test of haplotype concordance.

The N-test of haplotype concordance among siblings affected by some disease under investigation is used to decide whether there is a disease susceptibility gene linked to a marker locus or chromosomal region. The use of this test and appropriate modifications of it is briefly reviewed. The power of the ordinary N-test is then derived as a function of several parameters. The sample size needed to attain a given power is then derived. Some of the parameters are specified and the required sample sizes are given in tables for different values of the main unknown parameters.

Algorithms

Keys to the diagnosis of occult urologic disease in children.

All physicians who care for children should be aware of the many indications for further urologic examination. A straightforward algorithmic approach to urologic diagnosis is not possible. The physician must individualize and carefully weight the indications for the often-times expensive and uncomfortable tests that are required for urologic diagnosis. The reward is ample when a significant correctable lesion is recognized early enough for salvage on the basis of seemingly unrelated signs or symptoms.

Age Factors

A contig assembly program based on sensitive detection of fragment overlaps.

An effective computer program for assembling DNA fragments, the contig assembly program (CAP), has been developed. In the CAP program, a filter is used to eliminate quickly fragment pairs that could not possibly overlap, a dynamic programming algorithm is applied to compute the maximal-scoring overlapping alignment between each remaining pair of fragments, and a simple greedy approach is employed to assemble fragments in order of alignment scores. To identify the true fragment overlaps, the dynamic programming algorithm uses specially chosen sets of alignment parameters to tolerate sequencing errors and to penalize "mutational" changes between different copies of a repetitive sequence. The performance tests of the program on fragment data from genomic sequencing projects produced satisfactory results. The CAP program is efficient in computer time and memory; it took about 4 h to assemble a set of 1015 fragments into long contigs on a Sun workstation.

Algorithms

Diagnosing missed cases of spinal muscular atrophy in genome, exome, and panel sequencing data sets.

PURPOSE: We set out to develop a publicly available tool that could accurately diagnose spinal muscular atrophy (SMA) in exome, genome, or panel sequencing data sets aligned to a GRCh37, GRCh38, or T2T reference genome. METHODS: The SMA Finder algorithm detects the most common genetic causes of SMA by evaluating reads that overlap the c.840 position of the SMN1 and SMN2 paralogs. It uses these reads to determine whether an individual most likely has 0 functional copies of SMN1. RESULTS: We developed SMA Finder and evaluated it on 16,626 exomes and 3911 genomes from the Broad Institute Center for Mendelian Genomics, 1157 exomes and 8762 panel samples from Tartu University Hospital, and 198,868 exomes and 198,868 genomes from the UK Biobank. SMA Finder's false-positive rate was below 1 in 200,000 samples, its positive predictive value was greater than 96%, and its true-positive rate was 29 out of 29. Most of these SMA diagnoses had initially been clinically misdiagnosed as limb-girdle muscular dystrophy. CONCLUSION: Our extensive evaluation of SMA Finder on exome, genome, and panel sequencing samples found it to have nearly 100% accuracy and demonstrated its ability to reduce diagnostic delays, particularly in individuals with milder subtypes of SMA. Given this accuracy, the common misdiagnoses identified here, the widespread availability of clinical confirmatory testing for SMA, and the existence of treatment options, we propose that it is time to add SMN1 to the American College of Medical Genetics list of genes with reportable secondary findings after genome and exome sequencing.

Humans

Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.

MOTIVATION: The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. RESULTS: This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. AVAILABILITY AND IMPLEMENTATION: All associated code and documentation are available at https://github.com/farahpoor/smkhe.

Computer Security

Integrating Optical Genome Mapping into the Genetic Diagnostic Algorithm: Clinical Utility in Unresolved Autosomal Recessive Disorders from a Large Cohort.

INTRODUCTION: The identification of precise genetic etiologies is indispensable for the clinical management of monogenic disorders. However, conventional diagnostic methods and exome sequencing (ES) frequently fail to identify complex structural variations (SVs), leaving the genetic basis unexplained in approximately 30-60% of suspected cases. Optical genome mapping (OGM) emerges as a high-resolution technology capable of detecting cryptic SVs inaccessible to standard methodologies. METHODS: In this study, we evaluated the clinical utility of integrating OGM into the diagnostic algorithm for unresolved monogenic diseases. Following negative or inconclusive results from standard ES pipelines, OGM was applied to a targeted subset of patients (n = 7) selected from a comprehensive clinical cohort of 1,257 individuals with suspected genetic disorders. RESULTS: The integration of OGM identified candidate SVs that may represent the second allelic alteration in two distinct cases; however, confirmation through parental segregation analysis remains pending. Specifically, OGM identified an intronic insertion in the TTLL5 gene and a deletion in a putative regulatory region approximately 400 kb upstream of the NMNAT1 gene, both of which were missed by prior diagnostic testing. CONCLUSION: Our findings suggest that OGM has potential value in investigating the missing heritability of autosomal recessive disorders. By detecting candidate SVs invisible to conventional methods, OGM may warrant consideration as a complementary diagnostic approach following inconclusive ES; however, larger cohorts and confirmatory functional studies are needed to establish its clinical utility.

Autosomal recessive disorders

Cost-Effectiveness and the Economics of Genomic Testing and Molecularly Matched Therapies.

Cost-effectiveness analysis of precision oncology can help guide value-driven care. Next-generation sequencing is increasingly cost-efficient over single gene testing because diagnostic algorithms require multiple individual gene tests to determine biomarker status. Matched targeted therapy is often not cost-effective due to the high cost associated with drug treatment. However, genomic profiling can promote cost-effective care by identifying patients who are unlikely to benefit from therapy. Additional applications of genomic profiling such as universal testing for hereditary cancer syndromes and germline testing in patients with cancer may represent cost-effective approaches compared with traditional history-based diagnostic methods.

Humans

Stimulation of human peripheral blood lymphocytes with chironomid hemoglobin allergen (Chi t I).

Hemoglobins (Chi t I) of the dipteron species Chironomus thummi thummi are known to cause severe allergic diseases in humans. We tested the allergen-specific stimulation of human peripheral blood lymphocytes (PBL) by Chi t I and its nine main components. Further, we applied fragments of the well-analyzed component III, obtained by cleavage with trypsin as well as arginine protease. In this way, we screened the molecule in order to identify T-cell epitopes. The whole component was found to be immunogenic and to have regions demonstrating varying PBL stimulation. In addition, interindividual patterns of reactivity, probably due to genetic restriction, were found. A T-cell epitope could be shown to be within the site 98-111, as predicted by application of Rothbard's algorithms.

Allergens