Search PubMedSearch

SEARCH · Search PubMed

Results for “Genetic testing algorithm”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

125 records · Page 7Linked to original sources

Zone equalisation normalisation for improved alignment of epigenetic signal.

MOTIVATION: High-throughput genomic technologies have transformed our understanding of biological systems, yet direct comparison and visualisation of these complex datasets remains challenging. Existing normalisation methods often fail to align genomic signal across samples due to sensitivity to sequencing depth differences and localised high-signal artefacts, leading to inconsistent replicate behaviour and increased downstream variability. RESULTS: We introduce Zone Equalisation Normalisation (ZEN), a novel approach designed to improve cross-sample signal alignment of genomic data. ZEN rescales genomic signal based on variance estimated within biologically enriched regions, reducing the influence of extreme outliers while preserving underlying biological structure. Using a diverse collection of data and our new genome-wide benchmarking approach, we reveal that ZEN improves biological and technical replicate alignment across the majority of tested conditions and experimental platforms. We further show that this improved signal comparability is associated with fewer differential accessibility calls between technical replicates and a more conservative set of biological differences. Together, these results demonstrate that ZEN provides a complementary framework to improve the accuracy and reliability of genomic data analysis and that normalisation choice can affect downstream analyses and biological interpretation. AVAILABILITY AND IMPLEMENTATION: ZEN is available as an open-source Python package via conda and PyPI. Source code, documentation, tutorials, and code to reproduce the analyses are available at https://github.com/Genome-Function-Initiative-Oxford/Zone-Equalisation-Normalisation and Zenodo (https://doi.org/10.5281/zenodo.21067751).

Epigenesis, Genetic

Episode clustering in phylogenetic networks.

MOTIVATION: The classical duplication episode clustering (EC) model introduced by Guigó et al. in the 1990s provides a foundational approach for inferring genomic duplication events crucial to understanding genome evolution. This model clusters single gene duplications from a collection of gene trees at locations in the species tree to minimize the total number of such locations, called duplication episodes. However, it does not capture reticulate evolutionary histories. RESULTS: Here, we introduce NetEC, a novel extension of this problem to phylogenetic networks. To solve NetEC, we first develop a polynomial-time dynamic programming (DP) algorithm for testing whether a given set of network nodes can serve as episode locations. We then propose a main inference algorithm that utilizes this DP component to optimize the episode count; while the feasibility test runs in polynomial time, the full optimization has exponential worst-case complexity, and an optional heuristic mode is provided for larger instances. We also propose an extended episode analysis procedure that identifies additional genomic duplication candidates below reticulation nodes, complementing the main algorithm by resolving potential upward clustering of duplications induced by reticulation. We evaluate our method on simulated data and on an empirical Pandanales dataset comprising over 29 000 gene trees, demonstrating exact and accurate inference of genomic duplication events even in the presence of multiple reticulations. AVAILABILITY AND IMPLEMENTATION: All experiments were conducted using the NetEC tool (https://github.com/ppgorecki/netec), with all input data, scripts, and parameter settings for reproduction available in the same repository.

Phylogeny

Cost-justification analysis of prenatal maternal serum alpha-feto protein screening.

The costs to an insurer of a 10-year maternal serum alpha-feto protein (MSAFP) screening program were subtracted from future medical care costs avoided by the insurer (benefits) to examine whether such a program would be cost-justified from the perspective of a managed health care system (i.e., result in net costs greater than or equal to 0). The analysis considered MSAFP screening for neural tube defects (NTDs) alone and then was repeated to consider screening for both NTDs and Down's syndrome. Using a 5% discount rate for future dollars, the costs to the insurer of a screening program for NTDs alone over 10 years exceeded costs avoided by $10.00 per person screened. Adding screening for Down's syndrome using the same MSAFP test increased the net cost by $22.00 to a total of--$32.00 per screenee. The estimate of the cost to the insurer was sensitive to assumptions regarding the costs of medical care avoided, the expense of MSAFP, the proportion of screened women requiring a genetic amniocentesis, and the cost of that procedure. The conclusion that screening would not result in a cost savings to the insurer was not changed by reasonable assumptions regarding 1) the appropriate discount rate; 2) the costs of MSAFP; 3) the costs of genetic amniocentesis; 4) the sensitivity of MSAFP; 5) the proportion of the population requiring genetic amniocentesis; and 6) the costs of 10 years of medical care for someone affected by Down's syndrome or an NTD. Other analyses suggested that screening for NTDs or Down's syndrome would be cost-justified when viewed from the perspective of society. The present work suggests this conclusion does not hold when the perspective of the insurer is taken because avoided costs of care realized by society exceed those realized by the insurer.

Adult

Plasma inflammatory proteome profiles identify MASLD among children with overweight or obesity.

BACKGROUND & AIMS: Pediatric metabolic dysfunction-associated steatotic liver disease (MASLD) is increasingly prevalent among children with overweight or obesity, yet its early diagnosis remains a major clinical challenge. This study aimed to identify circulating inflammatory proteins associated with MASLD and to develop a proteomic risk score (ProScore) to improve diagnostic accuracy. METHODS: In this cross-sectional study of 161 children (median age 8.5&#xa0;years) with overweight or obesity, MASLD was assessed by vibration-controlled transient elastography, with 42 cases identified. Plasma concentrations of 92 inflammation-related proteins were quantified using a high-throughput proximity extension assay. The ProScore was compared with eleven conventional anthropometric/metabolic indices (WHtR, METS-IR, SPISE, PNFI, VAI, LAP, TyG, TyG-ALT, TyG-WC, TyG-WHtR, and TyG-BMI) and a genetic risk score (GRS). Six machine learning algorithms were employed and diagnostic performance was assessed using area under the curve (AUC) with fivefold cross-validation. RESULTS: Fifteen proteins were significantly associated with MASLD. A six-protein panel (FGF-21, CDCP1, CD244, OPG, Flt3L, MCP-1) achieved the highest diagnostic accuracy (AUC&#x2009;=&#x2009;0.84), exceeding that of all conventional indices (AUC&#x2009;=&#x2009;0.65-0.78; all P&#x2009;<&#x2009;0.05). ProScore performance remained robust in school-based validation (AUC&#x2009;=&#x2009;0.83), with no substantial improvement when combined with conventional indices. Diagnostic accuracy was higher in children with lower GRS (AUC&#x2009;=&#x2009;0.92) than in those with higher GRS (AUC&#x2009;=&#x2009;0.80; P&#x2009;=&#x2009;0.003). CONCLUSIONS: A proteomic signature of systemic inflammation provides accurate, non-invasive identification of MASLD in at-risk children, outperforming conventional metabolic and genetic tools, and may have utility in clinical and public health settings.

Humans

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans

A computer simulation model for analysis of conformation of nuclear chromatin and of the transcription process.

Based on experimental data (see Lindigkeit et al., 1974) an algorithmic computer model was devised with the following parameters: (a) nbr. and position of hypothetic blocker sites at the DNA which inhibit transcription in such a manner that distinct RNA chain lengths arise; (2) time depending probability functions that a polymerase molecule (p.m.) is able to pass a blocker site; (3) time depending rate of viability of synthetizing p.m. and (4) distribution of the p.m. on the template at the time t0. For comparison between observed and computed results 3 parameters are used: (1) distribution of chain lengths of RNA molecules in vitro synthetized; (2) their total amount and (3) the number of active p.m. By means of comparisons between computed and experimental results it is possible to test hypotheses about internal structural and functional parameters of the system under investigation, e.g. estimation of the rel. influence of template vs. p.m. characters in the transcription process, hypotheses about the type of distribution of p.m. at time t0, their initiation and salt depending removal probability of the blocker structures.

Cell Nucleus

Influence of aberrant observations on high-resolution linkage analysis outcomes.

Because of the availability of efficient, user-friendly computer analysis programs, the construction of multilocus human genetic maps has become commonplace. At the level of resolution at which most of these maps have been developed, the methods have proved to be robust. This may not be true in the construction of high-resolution linkage maps (3-cM interlocus resolution or less). High-resolution meiotic maps, by definition, have a low probability of recombination occurring in an interval. As such, even low frequencies of errors in typing (1.5% or less) may influence mapping outcomes. To investigate the influence of aberrant observations on high-resolution maps, a Monte Carlo simulation analysis of multipoint linkage data was performed. Introduction of error was observed to reduce power to discriminate orders, dramatically inflate map length, and provide significant support for incorrect over correct orders. These results appear to be due to the misclassification of nonrecombinant gametes as multiple recombinants. Chi 2-Like goodness-of-fit analysis appears to be quite sensitive to the appearance of misclassified gametes, providing a simple test for aberrant data sets. Multiple pairwise likelihood analysis appears to be less sensitive than does multipoint analysis and may serve as a check for map validity.

Algorithms

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs.

MOTIVATION: Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. RESULTS: We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635&#xa0;969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. AVAILABILITY AND IMPLEMENTATION: Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Genome-Wide Association Study

Relationship between Y-chromosome length and first-trimester spontaneous abortions.

The hypothesis that variation in Y-chromosome length is associated with repetitive fetal wastage was tested. Chromosome lengths were objectively quantitated by scanning photographic negatives of metaphases with a computer programmed to (1) select boundary thresholds and (2) construct and measure centerlines with a cubic spline-fitting algorithm. Variation in Y length among cells of different individuals was standardized by use of the ratio of the length of the Y to the average of the lengths of the No. 20s (20) in the same cell. Three groups were studied: (1) men whose wives had three or more spontaneous abortions and no live-born infants, (2) men whose wives had both abortions and normal live-born infants, and (3) control men whose wives had normal live-born infants only. Although central tendencies were similar in the three groups, the distributions of Y lengths among the three groups were significantly different (chi 2(6) = 15.33, 0.025 greater than p greater than 0.010). This difference was primarily because more of the subjects with only repetitive loss had Y lengths in the "tails" of the distribution rather than in the center. Our observations suggest the existence of an optimal Y length with respect to reproductive performance.

Abortion, Habitual

Site-directed deletion mutagenesis within the T4 endonuclease V gene: dispensable sequences within putative loop regions.

Endonuclease V from bacteriophage T4 may be one of the first DNA-repair enzymes to have its three-dimensional structure determined by X-ray crystallography (Morikawa et al., 1988). However, since this structure is not yet available, analyses of the sequence of the protein were performed in order to guide site-directed mutational studies of enzyme structure-function relationships. The enzyme is predominantly alpha-helical, so that an algorithm which finds the locations of turns or loops in the structure would be expected to approximately locate the helices along the sequence. Two loop sites were identified which might be adjacent in the tertiary structure according to a model developed from the loop predictions and the derived secondary structure. Deletion of three residues at each loop site produced protein molecules which retained considerable in vitro enzyme activity and in vivo repair function. However, the mutant proteins did not accumulate as well within the cell as the wild-type enzyme, suggesting that the nascent molecules folded inefficiently. Combination of the two deletions yielded a molecule with activity enhanced over one of the individual mutants, a result which can be interpreted as a classic second-site mutational reversion. This result supports the hypothesis that these regions are adjacent in the enzyme tertiary structure.

Base Sequence

Lineage-associated small inversions disrupt dosT, dnaE2, and a promoter-adjacent region in some Mycobacterium tuberculosis isolates.

UNLABELLED: Large molecular inversions in the genome of Mycobacterium tuberculosis (Mtb) due to factors like the presence of insertion sequences and transposases are widely known. However, smaller inversions within coding sequences and non-coding control elements are rarely reported. The present study aims to identify inversions and their potential impact on Mtb biology in a lineage-specific manner. Structural variants (SVs) could only be detected by long reads. For this, we simulated long reads by de novo assembling the short-read sequencing data sets and subsequently aligned representative strains from each lineage using the Progressive Mauve algorithm. Independently, long-read sequencing from the Pacific Biosciences platform was acquired and analyzed using the structural variant identification method. Variants were merged, and Fisher's exact test was carried out to identify the inversion association with lineages. To visualize deoxyribonucleic acid (DNA) features, the DNA-features-viewer tool was used. Simulated reads from short-read sequencing gave indications of lineage (L)-specific inversions. The long-read sequencing approach led to the identification of seven unique inversions: two positively associated with L1, one positively associated with L3, two negatively associated with L4, and two positively associated with L3 but negatively associated with L4 (P < 0.05). The inversions encompassed primarily non-essential genes like sdaA, dosT, Rv2026c, dnaE2, Rv1341, Rv1342, and lprD. An interesting inversion was observed in the upstream control element of purB and Rv0776c. The study sheds light on small inversions that may be causing alterations in expression, formation of fusion genes, and nonsense mutations that may have a role in lineage-specific phenotypic changes. IMPORTANCE: The role of mutations like SNPs and INDELs and their association with drug resistance is well known in Mycobacterium tuberculosis (Mtb). However, structural variations, especially inversions, are largely overlooked and unreported. In this paper, publicly available whole-genome sequencing datasets from Illumina and Pacific Biosciences-Oxford Nanopore Technologies platform have been used to detect inversions and report seven unreported Mtb lineage-specific small inversions.

Mycobacterium tuberculosis

Detecting Introgression in Shallow Phylogenies: How Minor Molecular Clock Deviations Lead to Major Inference Errors.

Recent theoretical and algorithmic advances in introgression detection, coupled with the growing availability of genome-scale data, have highlighted the widespread occurrence of interspecific gene flow across the tree of life. However, current methods largely depend on the molecular clock assumption-a questionable premise given empirical evidence of substitution rate variation across lineages. While such rate heterogeneity is known to compromise gene flow detection among divergent lineages, its impact on closely related taxa at shallow evolutionary timescales remains poorly understood, likely because these taxa are often assumed to adhere to a molecular clock. To address this gap, we combine theoretical analyses and simulations to evaluate the robustness of widely used site pattern methods (D-statistic and HyDe) to rate variation across phylogenetic timescales. Our results demonstrate that both methods exhibit high sensitivity to even minor deviations from the molecular clock at shallow timescales, complementing previous findings at deeper scales. Specifically, in young phylogenies (with an age of 3 &#xd7; 105 generations) with small population sizes, weak (17% difference) and moderate (33% difference) rate variation can inflate false-positive rates up to 35% and 100%, respectively, using site pattern counts from a 500&#x2005;Mb genome. Employing a more distant outgroup intensifies these spurious signals. Our study demonstrates that summary tests for introgression are pervasively vulnerable to minor rate variations and underscores the critical need for advanced methodologies to disentangle genuine introgression from false signals generated by rate heterogeneity.

Phylogeny

The distribution of the frequency of occurrence of nucleotide subsequences, based on their overlap capability.

DNA's genetic code can be represented as an alphabetic sequence composed of the four letters A, C, G, and T, which represent the four types of nucleotides--adenylic, cytidylic, guanylic, and thymidylic acid--of which DNA is composed. Now that these sequences have been identified for many genes and are available in computer-readable form, scientists can analyze these data and search for patterns in an attempt to learn more about the regulatory functions of the gene. One area of study is that of the frequency of occurrence of specific nucleotide subsequences (e.g., ACAC) within part or all of a nucleotide sequence. This paper derives the probability distribution of the frequency of occurrence of a subsequence within a nucleotide sequence, under the hypothesis that the four nucleotides occur at random and with equal probability. This distribution is nontrivial because different subsequences have different "overlap capability." For example, the subsequence AAAA can occur up to 17 times in a sequence of length 20 (which would happen if the sequence were composed solely of A's), but the subsequence ACGT cannot occur more than 5 times in a sequence of length 20. Thus, the frequency distributions are different for each type of overlap capability. It is of interest to assess and compare the degree of nonrandomness for different subsequences or among different portions of a sequence; the existence and degree of nonrandomness may be related to the type and degree of functionality of a nucleotide (sub)sequence. The frequency distributions provided here can be used to perform exact significance tests of the hypothesis of randomness. An approximate test is also described for use with long sequences; this can be used to test a more general null hypothesis of nucleotides occurring with unequal probabilities.

Algorithms

Inclusion of Multi-Omic Biomarkers Improves Prediction Accuracy of Response, Relapse, and Overall Survival in Acute Myeloid Leukemia Patients Receiving High-Intensity Induction Chemotherapy.

BACKGROUND: Despite advancements in genetic markers for acute myeloid leukemia (AML) risk stratification, outcome prediction remains challenging due to disease heterogeneity and dynamic genetic changes, highlighting the need for reliable biomarkers to improve AML treatment strategies and patient outcomes. To refine outcome predictions, we investigated the use of microbial-derived biomarkers to predict composite complete remission (CRc), relapse, and survival for patients on high- and low-intensity regimens, and to integrate those variables into the widely clinically utilized European Leukemia Network (ELN-2022) genetic risk classification model for high-intensity-treated patients. METHODS: We first developed machine learning models that integrate baseline fecal metabolomics, 16S rRNA-based stool microbiome features, and clinical metadata (sex, antibiotic administration, AML somatic mutations, and cytogenetics) from two cohorts of AML patients (n&#x2009;=&#x2009;83) undergoing remission induction chemotherapy. Univariate tests and sparse canonical correlation analysis were employed for variable selection and to explore fecal metabolite-microbe relationships. A robust machine learning approach using XGBoost was employed, with 100 stratified data splits (80% training, 20% testing) and coarse-to-fine hyperparameter optimization. Variable importance was aggregated across all models to select key predictors. RESULTS: For high-intensity-treated patients, XGBoost models achieved aggregated AUROC scores of 0.719, 0.729, and 0.65 for CRc, relapse, and overall survival, respectively. For low-intensity-treated patients, these models achieved aggregate AUROC scores of 0.945, 0.724, and 0.768 for these same outcomes, respectively. Integrating the biomarkers identified in the high-intensity machine-learning models with the current ELN-2022 AML risk stratification system effectively stratified patients into risk categories, which obtained higher concordance indices and likelihood ratios, demonstrating improved prognostic accuracy for each outcome compared to ELN-2022 alone. CONCLUSIONS: The inclusion of microbial-derived biomarkers serves as a robust prognostic tool to improve outcome prediction in AML patients, highlighting the potential of its integration into AML risk assessment and paving the way for personalized treatment strategies and improved patient outcomes.

Humans

Medical Research Council European trial of chorion villus sampling. MRC working party on the evaluation pf chorion villus sampling.

First-trimester chorion villus sampling has the advantage over second-trimester amniocentesis of allowing earlier prenatal diagnosis of various genetic and cytogenetic disorders in the fetus (and therefore earlier termination in affected pregnancies) but the relative safety and diagnostic accuracy remain unclear. Between 1985 and 1989, 3248 women seeking prenatal diagnosis, principally because of their age, were recruited to an international, multicentre, randomised comparison of the safety and diagnostic accuracy of the two techniques--5% of women allocated chorion villus sampling and 8% of those allocated amniocentesis were not tested, usually because of spontaneous miscarriage. 6% and 2% were retested, in most because of sampling failure. The endpoint of a liveborn infant who survived was achieved by 86% of women allocated chorion villus sampling and 91% of those allocated amniocentesis; statistical analysis, after appropriate weighting for a centre's contribution, showed that the typical difference between the groups was 4.6% (95% confidence interval 1.6-7.5%; p less than 0.01). This difference reflected more spontaneous fetal deaths before 28 weeks' gestation (2.9% [0.6-5.3%]); more terminations of pregnancy for chromosomal anomalies (1.0% [0.0-2.1%]); and more neonatal deaths (0.3% [-0.1 to 0.7%]). The difference in neonatal deaths was due to a preponderance of very immature liveborn infants in the chorion villus sampling group, and this factor also explained that group's longer mean stay in hospital. More abnormal diagnoses followed chorion villus than amniotic fluid analyses (5.6% vs 3.9%). This difference was largely due to diagnoses of trisomy 18 and of (usually mosaic) abnormalities known to be confined to the placenta. 3 terminated pregnancies were false positives, 1 tested by chorion villus sampling and 2 by amniocentesis, and 2 other mosaic cases diagnosed by chorion villus sampling may have been false positives. There was 1 false-negative result in the chorion villus sampling group. The possibility of earlier exclusion or diagnosis of some fetal disorders afforded by first-trimester chorion villus sampling must be set against its clinical risks.

Abortion, Induced

Factor IXHollywood: substitution of Pro55 by Ala in the first epidermal growth factor-like domain.

Factor IX is a multidomain protein essential for hemostasis. We describe a mutation in a patient affecting the first epidermal growth factor (EGF)-like domain of the protein. All exons and the promoter region of the gene were amplified by the polymerase chain reaction method, and sequenced. Only a single mutation (C----G), that predicts the substitution of Pro55 by Ala in the first EGF domain was found in the patient's gene. This mutation leads to new restriction sites for four enzymes. One new site (Nsi) was tested in the amplified exon IV fragment and was shown to provide a rapid and reliable marker for carrier detection and prenatal diagnosis in the affected family. The factor IX protein, termed factor IXHollywood (IXHW), was isolated to homogeneity from the patient's plasma. As compared with normal factor IX (IXN), IXHW contained the same amount of gamma-carboxy glutamic acid but twice the amount of beta-OH aspartic acid. Both IXHW and IXN contained no detectable free -SH groups. Further, IXHW could be readily cleaved to yield a factor IXa-like molecule by factor Xla/Ca2+. However, IXaHW (compared with IXaN) activated factor X approximately twofold slower in the presence of Ca2+ and phospholipid (PL), and 8- to 12-fold slower in the presence of Ca2+, PL, and factor VIIIa. Additionally, IXaHW had only approximately 10% of the activity of IXaN in an aPTT assay. In agreement with the nuclear magnetic resonance-derived structure of EGF, the Chou-Fasman algorithm strongly predicted a beta turn involving residues Asn-Pro55-Cys-Leu in IXN. Replacement of Pro55 by Ala gave a fourfold decrease in the beta turn probability for this peptide, suggesting a change(s) in the secondary structure in the EGF domain of IXHW. Since this domain of IXN is thought to have one high-affinity Ca2+ binding site and may be involved in PL and/or factor VIIIa binding, the localized secondary structural changes in IXHW could lead to distortion of the binding site(s) for the cofactor(s) and, thus, a dysfunctional molecule.

Alanine

[Investigation of algorithm for the calculation of probability of paternity likelihood using personal computer program, including the application to parentage testing in the decreased party].

Algorithm for the computerized calculation of probability of paternity likelihood was investigated. The probability is calculated by Essen-Möller's formula as W = X/(X+Y) = 1/(1 + Y/X). The X value is also given as X = Hl, m, n/Kl, m, where kl, m and Hl, m, n are the probabilities of mother-child and mother-child-father combinations, respectively. In this study, four functions as F(PQ) = [1-(P not equal to Q)] x p x q, Z(RS) = (1- (R = S)), K(PQ,RS) = 1/2([(R = P) + (R = Q)].s + [(S = P) + (S = Q)].r).F (PQ)/Z(RS) and H(PQ, RS, TU) = 1/4([(R = P) + (R = Q)] [(S = T) + (S = U)] + [(S = P) + (S = Q)] [(R = T) + (R = U)]) x F(PQ).F(TU)/Z(RS) were created, where PQ, RS and TU were the genotypes of mother, child and the alleged father, P, Q, R, S, T and U were their alleles, and p, q, r, s, t and u were the allele frequencies. The equality or inequality in parenthesis was the relation operator which gave -1 or 0 when the expression was true of false, respectively. Then, three formulae as Y = n sigma k = l F([TU]k), Kl, m = 1 sigma i = l m sigma j = l K ([PQ]i, [RS]j) and Hl, m, n = l sigma i = l m sigma j = l n sigma k = l H([PQ]i, [RS]j, [TU]k) were obtained, where [PQ]i, [RS]j and [TU]k were one of the mother's, one of the child's and one of the putative father's genotypes considered from their phenotypes, respectively. Using these formulae, the probability of paternity likelihood could be calculated in every case. These formulae were programmed in BASIC language using a personal computer. Algorithm for the calculation of the probability in the deceased party was also investigated.

Algorithms