Search PubMedSearch

SEARCH · Search PubMed

Results for “Genetic testing algorithm”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding

Search for faster methods of fitting the regressive models to quantitative traits.

The regressive models describe familial patterns of dependence of quantitative measures by specifying regression relationships among a person's phenotype and genotype and the phenotypes and genotypes of antecedents. When the number of sibs in the pattern of dependence increases, as in the class D regressive model, computation of the likelihood becomes time consuming, since the Elston-Stewart algorithm cannot be used generally. On the other hand, the simpler class A regressive model, which imposes a restriction on the sib-sib correlation, may lead to inference of a spurious major gene, as already observed in some instances. A simulation study is performed to explore the robustness of class A model with respect to false inference of a major gene and to search for faster methods of computing the likelihood under class D model. The class A model is not robust against the presence of a sib-sib correlation exceeding that specified by the model, unless tests on transmission probabilities are performed carefully: false detection of a major gene is reduced from a number of 26-30 to between 0 and 4 data sets out of 30 replicates after testing both the Mendelian transmission and the absence of transmission of a major effect against the general transmission model. Among various approximations of the likelihood formulation of the class D model, approximations 6 and 8 are found to work appropriately in terms of both the estimation of all parameters and hypothesis testing, for each generating model. These approximations lessen the computer time by allowing use of the Elston-Stewart algorithm.

Computer Simulation

Finlay-Wilkinson random regression for yield and yield stability prediction in cereals.

Year-to-year climate variability poses a challenge for agriculture by increasing crop yield variability; therefore, there is a need to identify genotypes that can withstand these fluctuations. With the right selection criteria, genotypes with yield stability across variable environmental conditions can be selected. Methods such as Finlay-Wilkinson random regression (FWRR) may allow us to use sparse datasets-common in plant breeding pipelines-and incorporate genomic data to leverage phenotypic information from related genotypes to predict yield stability. Our objective was to examine how the number of environments and the variance among those environments affect stability predictions. We also integrate FWRR as a genomic prediction tool for characterizing yield stability, comparing it to the traditional genomic prediction models as a reference. We used three datasets: one highly unbalanced dataset for oats (Avena sativa L.) and two completely balanced datasets with different numbers of environments for barley (Hordeum vulgare L.) and wheat (Triticum aestivum L.). We fit standard Finlay-Wilkinson (FW) and FWRR models to estimate grain yield and stability under various scenarios. We found that the estimated stability values obtained were similar using balanced datasets for FW or FWRR. FWRR also achieved moderate predictive ability for stability using unbalanced datasets under 10-fold cross-validation (CV1) with new genotypes. In terms of environmental representation, selecting the right set of environments for inclusion in the model was more important than adding more environments. Our results suggest the possibility of using FWRR to select stable genotypes earlier in line development, as well as to design resource-efficient stability-testing schemes.

Hordeum

A corpus of GA4GH phenopackets: Case-level phenotyping for genomic diagnostics and discovery.

The Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema was released in 2022 and approved by ISO as a standard for sharing clinical and genomic information about an individual, including phenotypic descriptions, numerical measurements, genetic information, diagnoses, and treatments. A phenopacket can be used as an input file for software that supports phenotype-driven genomic diagnostics and for algorithms that facilitate patient classification and stratification for identifying new diseases and treatments. There has been a great need for a collection of phenopackets to test software pipelines and algorithms. Here, we present Phenopacket Store. Phenopacket Store v.0.1.19 includes 6,668 phenopackets representing 475 Mendelian and chromosomal diseases associated with 423 genes and 3,834 unique pathogenic alleles curated from 959 different publications. This represents the first large-scale collection of case-level, standardized phenotypic information derived from case reports in the literature with detailed descriptions of the clinical data and will be useful for many purposes, including the development and testing of software for prioritizing genes and diseases in diagnostic genomics, machine learning analysis of clinical phenotype data, patient stratification, and genotype-phenotype correlations. This corpus also provides best-practice examples for curating literature-derived data using the GA4GH Phenopacket Schema.

Humans

Multivariate genetic evaluation in swine combining data from different testing schemes.

A computational strategy is presented that allows rapid implementation of genetic evaluations using multivariate mixed models. Data generated in different testing programs such as field tests of boars and gilts, litter recording schemes and station tests of sibs may be combined to provide an estimate of the aggregate genotype. Residual and additive genetic covariance structures are given for the multivariate evaluation of individual measurements and group averages because they often are collected for sib groups at test stations using a modified animal model. Pseudo code is given for the implementation of a "generic" testing structure illustrated by a numerical example based on six traits from field tests of boars and four traits from station tests of sibs. BLUPs for all six traits are calculated for boars, parents and sib groups. Aggregate genotypes that correspond to the selection indices commonly used are calculated for selection candidates.

Algorithms

A funny thing happened to us on the way to the latent entities.

Inferred latent entities, whether those of psychoanalysis, factor analysis, or cluster analysis, have declined in value for many clinical psychologists, both as tools of practice and as objects of theoretical interest. Behavior modification, rational-emotive therapy, crisis intervention, psycho-pharmacology, and actuarial prediction all tend to minimize reliance on latent entities in favor of purely dispositional concepts. Behavior genetics is, however, a powerful movement to the contrary. As regards categorical entities (types, taxa, syndromes, diseases), history reveals no impressive examples of their discovery by cluster algorithms; whereas organic medicine and psychopathology have both discovered many taxonic entities without reliance on formal (statistical) cluster methods. I offer eight reasons for this strange condition, with associated suggestions for ameliorating it. Adopting a realist instead of a fictionist approach to taxonomy, I give high priority to theory-based mathematical derivation of quantitative consistency tests for all taxometric results. I urge a large scale cooperative survey of taxometric methods based on Monte Carlo runs, biological pseudoproblems where the true axon is independently known, and live problem in genetics, organic medicine, and psychopathology. An empirical example of taxometric bootstrapping and consistency testing was presented from my own current research on schizotypy.

Humans

Crossover counts and likelihood in multipoint linkage analysis.

For large numbers of genetic loci, jointly tested to determine their order along a chromosome, likelihood methods become unfeasible owing to the very large numbers of discrete alternative hypotheses (locus orderings) whose likelihoods must be separately evaluated. A method to order loci according to the criterion of minimizing the obligatory crossover count is therefore proposed. A branch-and-bound algorithm implementing this proposal has been programmed; the properties of this algorithm are investigated. The statistical properties of the proposed method are also considered. It is shown to be consistent under wide conditions, including arbitrary locus spacings, variable amounts of information per locus, and some patterns of interference. The relationship between the minimum crossover order and the maximum likelihood order is discussed. For fully informative gametes, and tight linkage, there is a virtual equivalence of the two criteria. For looser linkage, there remains a close relationship.

Algorithms

Bent helical structure in Trypanosoma lewisi minicircles.

By gel retardation assay and computational analysis we demonstrated a bent region in Trypanosoma lewisi, localized in two different classes of minicircles. We showed that in each minicircle this bent region is unique, adjacent to one of two highly conserved regions and characterized by adenine stretches. The same properties are conserved in the majority of minicircles from Trypanosomes tested so far. Therefore, we suggest that the genetic information could be located in a definite structure of minicircle DNA molecules rather than in the nucleotide sequence.

Adenine

RAREsim2: flexible simulation of rare variant genetic data using real haplotypes.

MOTIVATION: Realistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., case-control status, technological or batch effects) and variant-level differences to represent a variety of causal models. RESULTS: We demonstrate RAREsim2's utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios, including various genetic ancestries, gene sizes, strengths of association, and proportions of risk variants. Type I Error was maintained and the test with the highest power matched previously known patterns. Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. AVAILABILITY AND IMPLEMENTATION: The RAREsim2 Python package is publicly available on Github (https://github.com/Hendricks-Research-Team/RAREsim2), PyPI (https://pypi.org/project/raresim/), and Zenodo (https://doi.org/10.5281/zenodo.19442523). Code for the example demonstration can be found at https://github.com/JessMurphy/RAREsim2-demo.

Software

Robust methods for the detection of genetic linkage for quantitative data from pedigrees.

The robust method for detecting linkage developed by Haseman and Elston [The investigation of linkage between a quantitative trait and a marker locus. Behav Genet 2:3-19, 1972] for data from sib pairs is extended to any type of noninbred relative pair. The regression of the squared relative-pair trait difference on the estimated proportion of genes identical by descent (i.b.d.) at a marker locus is shown to depend upon the recombination fraction between the two loci; the regression coefficient is negative if the trait and marker loci are linked. A test for linkage based on data from any informative type of relative pair can thus be obtained by testing that this regression coefficient is less than zero. Formulae for the asymptotic power of such tests for linkage based upon independent relative pairs are developed. Results are also given for the special case in which the proportion of genes shared i.b.d. for relative pairs is known. Finally, a general algorithm is described that will incorporate all available pedigree data to calculate an estimate of the proportion of genes that a relative pair shares i.b.d. at a marker locus.

Genetic Linkage

Determination of near-optimum use of hospital diagnostic resources using the "GENES" genetic algorithm shell.

"GENES", a genetic algorithm shell developed by the authors, was used to optimize allocation of hospital resources for a small set of hypothetical patients. GENES creates a random population of rule sets of the IF..THEN type, which are variable in both the number of rules in each set and in the size of each rule. GENES applies each rule set to a patient data base, ranks the goodness of each set as applied, and uses the mechanisms of population genetics, i.e. mutation, crossover, inversion and survival of the fittest, to create a new, and often improved generation of rule sets. It also allows for the time dependent nature of medical tests, the possibility of injury associated with those tests, and the fact that results may not always be conclusive. Using 10-11 artificially created patients admitted under the diagnosis of possible gall bladder disease, a rule set was obtained which selected a testing strategy from a list of available hospital resources and correctly diagnosed all patients at minimum cost in no more than 3807 generations.

Algorithms

Incidence patterns and genetic validation of primary glaucoma subtypes among 1 million adults in China and the UK.

BACKGROUND/AIMS: Primary open-angle glaucoma (POAG) and primary angle-closure glaucoma (PACG) are distinct diseases, yet many glaucoma cases in population-based datasets lack subtype specification. We assessed incidence patterns of glaucoma subtypes in China and the UK and used genetic evidence to infer the likely subtype composition of cases recorded as unspecified glaucoma. METHODS: Incident primary glaucoma was identified from linked inpatient records in the prospective China Kadoorie Biobank (CKB; n=512 504) and UK Biobank (UKB; n=492 329) studies. Cohort-specific phenotyping algorithms defined POAG, PACG and unspecified glaucoma. Adjusted incidence rates were estimated by direct standardisation. To support subtype inference, polygenic risk scores (PRSs) were constructed using ancestry-specific genome-wide association studies, including a new East Asian PACG meta-analysis, and tested for association with glaucoma phenotypes using multivariable logistic regression. RESULTS: Over 12 years of follow-up, 1658 primary glaucoma cases were identified in CKB and 7643 in UKB. Most (>68%) cases lacked subtype specification. Incidence increased with age and was twofold higher among women for PACG in both cohorts and for unspecified glaucoma in CKB. In UKB, POAG incidence was fivefold higher among Black than White participants, with a similar but attenuated pattern for unspecified glaucoma. PRS analyses indicated that unspecified glaucoma closely aligned with PACG in CKB but was more heterogeneous in UKB. CONCLUSION: Healthcare-recorded incidence patterns for POAG and PACG were consistent with established demographic risk factors, whereas unspecified glaucoma showed differences in subtype composition between populations. Integrating epidemiological and genetic evidence improves interpretation of glaucoma phenotypes when detailed clinical information is unavailable.

Epidemiology

Finite-state models in the alignment of macromolecules.

Minimum message length encoding is a technique of inductive inference with theoretical and practical advantages. It allows the posterior odds-ratio of two theories or hypotheses to be calculated. Here it is applied to problems of aligning or relating two strings, in particular two biological macromolecules. We compare the r-theory, that the strings are related, with the null-theory, that they are not related. If they are related, the probabilities of the various alignments can be calculated. This is done for one-, three-, and five-state models of relation or mutation. These correspond to linear and piecewise linear cost functions on runs of insertions and deletions. We describe how to estimate parameters of a model. The validity of a model is itself an hypothesis and can be objectively tested. This is done on real DNA strings and on artificial data. The tests on artificial data indicate limits on what can be inferred in various situations. The tests on real DNA support either the three- or five-state models over the one-state model. Finally, a fast, approximate minimum message length string comparison algorithm is described.

Algorithms

Blood-based epigenome-wide analyses of chronic low-grade inflammation across diverse population cohorts.

Chronic inflammation is a hallmark of age-related disease states. The effectiveness of inflammatory proteins including C-reactive protein (CRP) in assessing long-term inflammation is hindered by their phasic nature. DNA methylation (DNAm) signatures of CRP may act as more reliable markers of chronic inflammation. We show that inter-individual differences in DNAm capture 50% of the variance in circulating CRP (N = 17,936, Generation Scotland). We develop a series of DNAm predictors of CRP using state-of-the-art algorithms. An elastic-net-regression-based predictor outperformed competing methods and explained 18% of phenotypic variance in the Lothian Birth Cohort of 1936 (LBC1936) cohort, doubling that of existing DNAm predictors. DNAm predictors performed comparably in four additional test cohorts (Avon Longitudinal Study of Parents and Children, Health for Life in Singapore, Southall and Brent Revisited, and LBC1921), including for individuals of diverse genetic ancestry and different age groups. The best-performing predictor surpassed assay-measured CRP and a genetic score in its associations with 26 health outcomes. Our findings forge new avenues for assessing chronic low-grade inflammation in diverse populations.

Humans

Expectation maximization algorithm for identifying protein-binding sites with variable lengths from unaligned DNA fragments.

An Expectation Maximization algorithm for identification of DNA binding sites is presented. The approach predicts the location of binding regions while allowing variable length spacers within the sites. In addition to predicting the most likely spacer length for a set of DNA fragments, the method identifies individual sites that differ in spacer size. No alignment of DNA sequences is necessary. The method is illustrated by application to 231 Escherichia coli DNA fragments known to contain promoters with variable spacings between their consensus regions. Maximum-likelihood tests of the differences between the spacing classes indicate that the consensus regions of the spacing classes are not distinct. Further tests suggest that several positions within the spacing region may contribute to promoter specificity.

Algorithms

Identifying potential tRNA genes in genomic DNA sequences.

We have developed an algorithm that automatically and reproducibly identifies potential tRNA genes in genomic DNA sequences, and we present a general strategy for testing the sensitivity of such algorithms. This algorithm is useful for the flagging and characterization of long genomic sequences that have not been experimentally analyzed for identification of functional regions, and for the scanning of nucleotide sequence databases for errors in the sequences and the functional assignments associated with them. In an exhaustive scan of the GenBank database, 97.5% of the 744 known tRNA genes were correctly identified (true-positives), and 42 previously unidentified sequences were predicted to be tRNAs. A detailed analysis of these latter predictions reveals that 16 of the 42 are very similar to known tRNA genes, and we predict that they do, in fact, code for tRNA, yielding a false-positive rate for the algorithm of 0.003%. The new algorithm and testing strategy are a considerable improvement over any previously described strategies for recognizing tRNA genes, and they allow detections of genes (including introns) embedded in long genomic sequences.

Algorithms

Testing for bimodality in frequency distributions of data suggesting polymorphisms of drug metabolism--hypothesis testing.

1. The theory of methods of hypothesis testing in relation to the detection of bimodality in density distributions is discussed. 2. Practical problems arising from these methods are outlined. 3. The power of three methods of hypothesis testing was compared using simulated data from bimodal distributions with varying separation between components. None of the methods could determine bimodality until the separation between components was 2 standard deviation units and could only do so reliably (greater than 90%) when the separation was as great as 4-6 standard deviation units. 4. The robustness of a parametric and a non-parametric method of hypothesis testing was compared using simulated unimodal distributions known to deviate markedly from normality. Both methods had a high frequency of falsely indicating bimodality with distributions where the components had markedly differing variances. 5. A further test of robustness using power transformation of data from a normal distribution showed that the algorithms could accurately determine unimodality only when the skew of the distribution was in the range 0-1.45.

Computers

Diagnosis of twin zygosity by mailed questionnaire.

A deterministic questionnaire method for zygosity determination is developed for use in epidemiological studies of adult twins. It is based on the answers of both members of a twin pair to two questions on similarity and confusion in childhood. The algorithm of the method is used to determine the zygosity status of a twin pair at two different levels of certainty. The validity of the method is tested by making blood marker determinations of 11 polymorphic marker systems fro a random sample of 104 twin pairs. The agreement between questionnaire and blood marker diagnosis was 100%, but the stricter level of certainty left 8.7% in the nonclassified group. The genetical representativeness of the sample is tested by the allele distribution of the markers as compared to the Finnish population data as well as by the distribution of the number of intra-pair differences in blood markers.

Blood Group Antigens