Search PubMed⌕ Search

Biomedical subjects

Lisa J Strug

Publications and source records attributed to Lisa J Strug.

6 recordsLinked to original sources

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software↗

An alternative foundation for the planning and evaluation of linkage analysis. II. Implications for multiple test adjustments.

The 'multiple testing problem' currently bedevils the field of genetic epidemiology. Briefly stated, this problem arises with the performance of more than one statistical test and results in an increased probability of committing at least one Type I error. The accepted/conventional way of dealing with this problem is based on the classical Neyman-Pearson statistical paradigm and involves adjusting one's error probabilities. This adjustment is, however, problematic because in the process of doing that, one is also adjusting one's measure of evidence. Investigators have actually become wary of looking at their data, for fear of having to adjust the strength of the evidence they observed at a given locus on the genome every time they conduct an additional test. In a companion paper in this issue (Strug & Hodge I), we presented an alternative statistical paradigm, the 'evidential paradigm', to be used when planning and evaluating linkage studies. The evidential paradigm uses the lod score as the measure of evidence (as opposed to a p value), and provides new, alternatively defined error probabilities (alternative to Type I and Type II error rates). We showed how this paradigm separates or decouples the two concepts of error probabilities and strength of the evidence. In the current paper we apply the evidential paradigm to the multiple testing problem - specifically, multiple testing in the context of linkage analysis. We advocate using the lod score as the sole measure of the strength of evidence; we then derive the corresponding probabilities of being misled by the data under different multiple testing scenarios. We distinguish two situations: performing multiple tests of a single hypothesis, vs. performing a single test of multiple hypotheses. For the first situation the probability of being misled remains small regardless of the number of times one tests the single hypothesis, as we show. For the second situation, we provide a rigorous argument outlining how replication samples themselves (analyzed in conjunction with the original sample) constitute appropriate adjustments for conducting multiple hypothesis tests on a data set.

Chromosome Mapping↗

An alternative foundation for the planning and evaluation of linkage analysis. I. Decoupling "error probabilities" from "measures of evidence".

The lod score, which is based on the likelihood ratio (LR), is central to linkage analysis. Users interpret lods by translating them into standard statistical concepts such as p values, alpha levels and power. An alternative statistical paradigm, the Evidential Framework, in contrast, works directly with the LR. A key feature of this paradigm is that it decouples error probabilities from measures of evidence. We describe the philosophy behind and the operating characteristics of this paradigm--based on new, alternatively-defined error probabilities. We then apply this approach to linkage studies of a genetic trait for: I. fully informative gametes, II. double backcross sibling pairs, and III. nuclear families. We consider complete and incomplete penetrance for the disease model, as well as using an incorrect penetrance. We calculate the error probabilities (exactly for situations I and II, via simulation for III), over a range of recombination fractions, sample sizes, and linkage criteria. We show how to choose linkage criteria and plan linkage studies, such that the probabilities of being misled by the data (i.e. concluding either that there is strong evidence favouring linkage when there is no linkage, or that there is strong evidence against linkage when there is linkage) are low, and the probability of observing strong evidence in favour of the truth is high. We lay the groundwork for applying this paradigm in genetic studies and for understanding its implications for multiple tests.

Computer Simulation↗

Possible interaction between HLA-DRbeta1 and thyroglobulin variants in Graves' disease.

Graves' disease (GD) is influenced by two major susceptibility loci, HLA-DR3 and thyroglobulin (Tg). Recently we have shown that specific HLA-DR and Tg gene sequences predispose to Graves' disease. Individuals carrying at least one arginine at position 74 of the DRbeta1 chain (denoted the R- genotype) have a significantly increased risk of GD, as do individuals homozygous for the single nucleotide protein (SNP) in exon 33 of the Tg gene (denoted the CC genotype). Therefore, for the current study we hypothesized that these two genes may interact to influence the etiology of GD. To test this hypothesis, we analyzed the genotypes of 185 Caucasian patients with GD and 143 Caucasian controls for both genes. We tested for an interaction effect, that is, is one gene's effect on GD greater when the other gene is also present than when the other gene is absent? A logistic regression analysis yielded an estimate of 4.31 for the interaction term (p = 0.053). Our results may suggest an interaction between the R- and CC variants in conferring susceptibility to GD. These results, if confirmed, may imply that these two variants interact biologically to increase the odds of GD.

Adult↗

Construction of the model for the Genetic Analysis Workshop 14 simulated data: genotype-phenotype relationships, gene interaction, linkage, association, disequilibrium, and ascertainment effects for a complex phenotype.

The Genetic Analysis Workshop 14 simulated dataset was designed 1) To test the ability to find genes related to a complex disease (such as alcoholism). Such a disease may be given a variety of definitions by different investigators, have associated endophenotypes that are common in the general population, and is likely to be not one disease but a heterogeneous collection of clinically similar, but genetically distinct, entities. 2) To observe the effect on genetic analysis and gene discovery of a complex set of gene x gene interactions. 3) To allow comparison of microsatellite vs. large-scale single-nucleotide polymorphism (SNP) data. 4) To allow testing of association to identify the disease gene and the effect of moderate marker x marker linkage disequilibrium. 5) To observe the effect of different ascertainment/disease definition schemes on the analysis. Data was distributed in two forms. Data distributed to participants contained about 1,000 SNPs and 400 microsatellite markers. Internet-obtainable data consisted of a finer 10,000 SNP map, which also contained data on controls. While disease characteristics and parameters were constant, four "studies" used varying ascertainment schemes based on differing beliefs about disease characteristics. One of the studies contained multiplex two- and three-generation pedigrees with at least four affected members. The simulated disease was a psychiatric condition with many associated behaviors (endophenotypes), almost all of which were genetic in origin. The underlying disease model contained four major genes and two modifier genes. The four major genes interacted with each other to produce three different phenotypes, which were themselves heterogeneous. The population parameters were calibrated so that the major genes could be discovered by linkage analysis in most datasets. The association evidence was more difficult to calibrate but was designed to find statistically significant association in 50% of datasets. We also simulated some marker x marker linkage disequilibrium around some of the genes and also in areas without disease genes. We tried two different methods to simulate the linkage disequilibrium.

Computer Simulation↗

Disease severity in siblings with cystic fibrosis.

Since cross-infection occurs between cystic fibrosis (CF) siblings, we hypothesized that subsequent siblings may acquire respiratory pathogens at an earlier age and have a more severe course of pulmonary disease. We studied a retrospective cohort of 31 sibling pairs from the CF database at the Hospital for Sick Children. Kaplan-Meier curves and modified log-rank tests were used to test sibling differences in age of acquisition of Pseudomonas aeruginosa (PA), Staphylococcus aureus (SA), or any positive culture. Differences in disease severity outcomes were explored. Older siblings were more likely to have both SA and any CF pathogen first isolated from respiratory culture at an older age than younger siblings (P = 0.0050 and P = 0.0008, respectively, by modified log-rank tests). However, more of the older siblings were positive on first culture at time of diagnosis, introducing an age-of-diagnosis bias. Hospitalization rates, courses of oral antibiotics, FEV(1) % predicted, and weight and height measurements were not better in the older children. No differences in clinical parameters were found between older and younger siblings. The apparent finding of younger age at first isolation of pathogens from respiratory cultures in younger siblings is likely because many older siblings were already infected with these organisms at time of diagnosis.

Age of Onset↗