Search PubMedSearch

SEARCH · Search PubMed

Results for “methods”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Newly Developed Structure-Based Methods Do Not Outperform Standard Sequence-Based Methods for Large-Scale Phylogenomics.

Recent developments in protein structure prediction have allowed the use of this previously limited source of information at genome-wide scales. It has been proposed that the use of structural information may offer advantages over sequences in phylogenetic reconstruction, due to their slower rate of evolution and direct correlation to function. Here, we examined how recently developed methods for structure-based homology search and tree reconstruction compare with current state-of-the-art sequence-based methods in reconstructing genome-wide collections of gene phylogenies (i.e. phylomes). While structure-based methods can be useful in specific scenarios, we found that their current performance does not justify using the newly developed structure-based methods as a default choice in large-scale phylogenetic studies. On the one hand, the best performing sequence-based tree reconstruction methods still outperform structure-based methods for this task. On the other hand, structure-based homology detection methods provide larger lists of candidate homologs, as previously reported. However, this comes at the expense of missing hits identified by sequence-based methods, as well as providing sets of homolog candidates with higher fractions of false positives. These insights help to guide the use of structural data in comparative genomics and highlight the need to continue improving structure-based approaches. Our pipeline is fully reproducible and has been implemented in a Snakemake workflow. This will facilitate a continuous assessment of future improvements of structure-based tools in the AlphaFold era.

Phylogeny

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

A universal, high-quality, and high-yield DNA purification method for mycobacteria, including Mycobacterium tuberculosis: large-scale assessment of the chloroform-bead method.

UNLABELLED: Genomic analysis of mycobacteria has become increasingly crucial for understanding drug-resistance mechanisms, molecular epidemiology, and pathogenesis. However, efficient extraction of high-molecular-weight genomic DNA from these organisms remains challenging because of their thick mycolic acid-rich cell walls. In this study, we report the chloroform-bead method, a universal DNA extraction protocol that combines chemical and mechanical disruptions to overcome these challenges. Multi-laboratory evaluation (16 sites) demonstrated the chloroform-bead method's superiority over conventional methods for Mycobacterium tuberculosis (DNA yield: 17.9 vs 1.9 &#xb5;g, purity A260/A230: 1.86 vs 1.22, both P < 0.001). Single-facility assessment extended these findings to >32 nontuberculous mycobacterial species (n = 1,058), showing performance comparable to M. tuberculosis (n = 1,000), with both achieving median yields of 22.2 &#xb5;g DNA and consistent quality metrics. The chloroform-bead method significantly reduced the processing time from 2 to 3 days to 2 h while ensuring complete sample sterilization, eliminating the need for species-specific optimization. This streamlined and universally applicable protocol represents a practical advancement in mycobacterial DNA extraction methodology, ideal for high-throughput genomic studies and routine clinical diagnostics. IMPORTANCE: Mycobacterial genomics is crucial for understanding pathogenesis and drug resistance; however, DNA extraction remains a significant challenge because of its unique cell wall. Traditional methods rely on enzymatic treatments, resulting in complex and time-consuming protocols with variable results. The chloroform-bead method introduces a paradigm shift by chemically and mechanically disrupting the mycolic acid layer and eliminating the need for enzymatic treatment. This standardized approach ensures consistent, high-quality DNA extraction across diverse mycobacterial species, thereby enhancing research capabilities and clinical applications.

Chloroform

Mixed Methods Research on Family Caregiving for Stroke Survivors: A Methodological Systematic Review.

AIM: To examine how mixed methods research has been applied in studies of family caregiving for stroke survivors, focusing on key methodological components (rationale, design types, integration strategies, and use of joint displays). DESIGN: Methodological systematic review. METHODS: A systematic search of five databases yielded 17 studies. The extraction focused on mixed methods features (rationale, design, integration, joint displays), and quality was appraised using the Mixed Methods Appraisal Tool. DATA SOURCES: PubMed, CINAHL, Scopus, Web of Science, and PsycINFO were searched for relevant studies published from 2010 to 2025. RESULTS: The included studies addressed topics such as caregiver burden, coping, resilience, and intervention outcomes. Convergent and explanatory sequential designs predominated. Complementarity was the most frequent rationale for mixing methods. Integration occurred mainly through merging, with fewer instances of connecting or building. Three studies included joint displays to integrate the results. CONCLUSION: Mixed methods research is increasingly applied in family caregiving. To advance the field, researchers should strengthen integration during analysis and results and improve transparency in reporting key design features. IMPLICATIONS FOR THE PROFESSION AND/OR PATIENT CARE: Strengthening methodological rigour in mixed methods studies on stroke caregiving will improve the evidence base for nursing practice. Intentional and meaningful integration of qualitative and quantitative evidence can better inform effective interventions and support programs, ultimately enhancing care for stroke survivors and their families. IMPACT: This review evaluates how mixed methods research is applied in family caregiving studies. It identifies significant methodological gaps, including unclear reporting of design and limited use of advanced integration techniques. The recommendations provide practical guidance for researchers to improve reporting and integration, yielding richer evidence to inform interventions and policies that support family caregivers. REPORTING METHOD: The review followed the PRISMA 2021 guidelines for transparent reporting of systematic reviews. PATIENT OR PUBLIC CONTRIBUTION: No patient or public involvement.

Humans

Next-Generation Sequencing Methods for Sensitive Hepatitis B Viral Genome Analysis: A European Study.

This multicentre study investigated the utility of next-generation sequencing (NGS) to detect and generate hepatitis B virus (HBV) genomes in samples of low viral load (from 0.2 to 6207 IU/mL). 23 HBV DNA-positive plasma samples of genotypes A-E and one HBV-negative control sample were assayed blindly via 9 established NGS methods from 6 European laboratories. Methods included untargeted metagenomics, pre-enrichment by probe-capture followed by Illumina sequencing, and HBV-specific PCR pre-amplification followed by sequencing with Nanopore or Illumina. Full HBV genomes were obtained only from samples with viral loads >&#x2009;1000 IU/mL using probe-capture methods, >&#x2009;200 IU/mL using PCR-Illumina methods, >&#x2009;10 IU/mL using PCR-Nanopore methods, and in no samples using metagenomic methods. Contamination was observed in the negative control and samples with very low viral loads in PCR-based methods. Probe-capture and metagenomic methods detected additional viruses not routinely screened in blood donations, including polyomaviruses and herpesviruses; positive results were confirmed by PCR. In conclusion, NGS may delineate whole-genome sequences at low viral loads if supported by a PCR pre-amplification step. Probe-capture methods also reliably detect HBV without pre-amplification but show limited genome coverage for samples with low viral loads; they may additionally detect a wide range of blood-borne viruses.

Humans

Comparison of culture and culture-free methods for comprehensive identification of mycobacteria: a single-center prospective study.

The genus Mycobacterium, including Mycobacterium tuberculosis and over 200 nontuberculous mycobacteria (NTM), shows wide variability in clinical outcomes and drug susceptibility. Although culture-based identification remains the gold standard, slow mycobacterial growth delays diagnosis and treatment. In this study, we evaluated a novel culture-free method for subspecies-level identification directly from sputum. In this single-center prospective cohort study at Osaka Toneyama Medical Center, we analyzed 125 sputum samples from 115 patients with NTM pulmonary disease and 10 with non-NTM respiratory conditions. Samples were decontaminated using N-acetyl-L-cysteine-sodium hydroxide (NALC-NaOH) or succinic acid. We compared the reference culture method (mycobacterial culture plus whole-genome sequencing) and a culture-free direct target capture sequencing method. Core genome multi-locus sequence typing identified subspecies in both workflows, covering 186 mycobacterial species, including M. tuberculosis. The 115 NTM cohort specimens yielded 57 smear-positive and 93 culture-positive results. The identified subspecies included 48 Mycobacterium avium subsp. hominissuis, 22 Mycobacterium intracellulare subsp. intracellulare, 5 subsp. chimaera, 7 Mycobacterium abscessus subsp. abscessus, 5 subsp. massiliense, 1 M. tuberculosis, and 5 other NTM species. The culture-free method showed a high identification rate for smear-positive specimens (75.4%) but a low identification rate for smear-negative specimens (13.9%). NALC-NaOH pretreatment resulted in higher accuracy (90.5%) than did succinic acid pretreatment (66.7%). Thus, our culture-free subspecies-level identification method achieved high accuracy, especially in alkaline-treated smear-positive sputum samples, achieving rates above 90%. This method is recommended in clinical practice for patients who require rapid diagnosis and timely initiation of appropriate treatment, bypassing time-consuming culture steps.IMPORTANCEAccurate identification of Mycobacterium species and subspecies is crucial for effective treatment, as drug susceptibility and clinical outcomes vary significantly among them. However, conventional diagnosis relies on culture-based methods that can take several weeks, critically delaying appropriate therapy. This study validates a novel culture-free method using target capture sequencing for the comprehensive, subspecies-level identification of over 186 mycobacterial species directly from sputum specimens. Our findings revealed the high accuracy of this approach for smear-positive specimens, especially with alkaline pretreatment. This rapid method is applicable in clinical settings and enables timely and precise treatment decisions, greatly benefiting patients who require urgent intervention.

Humans

Disagreement-informed arbitration for gene regulatory network inference: A score-level meta-classifier and a diagnostic typology of inter-method conflict.

Gene regulatory network inference methods routinely disagree about individual edges, and practitioners resolve those conflicts by choosing one method or averaging them all. We ask whether the conflict can instead be arbitrated per edge. A gradient-boosted classifier is trained on the raw scores that ten inference methods-correlation-based, information-theoretic, sparse-regression and tree-ensemble, including GENIE3, GRNBoost2, CLR and ARACNe-assign to each candidate regulator-target pair, so that the weight given to each method varies from edge to edge. Across six single-cell perturbation screens spanning four cell types, arbitration improves on mean ensembling by +0.056 AUROC on Adamson and +0.083 on Shifrut under target-grouped cross-validation. The evaluation protocol turns out to matter more than the model. Edge-level cross-validation, standard in this literature, inflates apparent gains by 0.060 AUROC through target-gene leakage-comparable to the entire honest improvement. The effect is far larger for methods that represent genes implicitly: a supervised graph-attention link predictor trained on identical folds scores AUROC 0.930 under edge-level cross-validation, better than anything else we evaluate, and 0.533 once target genes are held out. Any method that parameterises genes is exposed, which covers most graph- and embedding-based approaches. A five-category typology of inter-method conflict localises where arbitration pays off, with the largest gains on edges where the methods disagree and the smallest where they already agree, while adding nothing as model input; we therefore report it as a diagnostic instrument rather than a modelling contribution. We also characterise what the ground truth measures: most perturbed genes in widely used screens are not transcription factors, and a mediation screen bounds how much of the perturbation response can be direct.

Ensemble methods

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results

CRISPR/Cas9-compatible plasmids enabling seven dominant genetic selection methods for the human fungal pathogen Cryptococcus neoformans.

Cryptococcus neoformans is the most common cause of human fungal meningitis and an important model system for studying fundamental eukaryotic biology. Genetic manipulation of this organism relies on three dominant drug resistance markers (nourseothricin acetyltransferase [NAT], neomycin phosphotransferase II [NEO], and hygromycin B phosphotransferase [HYG]) and the recyclable dominant prototrophic marker amdS. With ongoing technological advances that are expanding our ability to explore cryptococcal gene function, contemporary studies often require multiple genetic manipulations in the same strain. Additional dominant selection methods would maximize the utility of these tools by facilitating their combinatorial use. Here, we identify blasticidin S resistance via the blasticidin S deaminase (BSD) or blasticidin S resistance (BSR) markers as a novel dominant selection method for C. neoformans. We further validate phleomycin resistance via the bleomycin resistance gene (BLE) marker as an additional selection method, confirming a study that first established this marker 25 years ago (J. Hua, J. D. Meyer, and J. K. Lodge, Clin Diagn Lab Immunol 7:125-128, 2000, https://doi.org/10.1128/cdli.7.1.125-128.2000). To enable highly efficient CRISPR/Cas9-mediated genome modification, we incorporated these markers, as well as the newly established dominant prototrophic marker ptxD (M. Khongthongdam, T. Phetruen, and S. Chanarat, Microbiol Spectr 13:e01618-24, 2025, https://doi.org/10.1128/spectrum.01618-24), into a vector series that enables the construction of fused marker-sgRNA products via PCR. Altogether, this work expands the number of dominant genetic selection methods for C. neoformans to seven, including five drug selection regimes and two prototrophic methods. The vector series has been deposited at Addgene. IMPORTANCE Cryptococcus neoformans is the top-ranked World Health Organization priority fungal pathogen due to its widespread distribution and inadequate treatment options. Additionally, as a basidiomycete yeast occupying an underexplored branch of the fungal kingdom, this organism is a powerful system for deciphering core eukaryotic biology that is absent in classic model fungi. Defining functions for novel cryptococcal genes is a crucial priority, and the availability of additional genetic selection methods would facilitate these efforts. In this study, we establish blasticidin S resistance as a novel genetic selection method for C. neoformans, and we validate a previous report using phleomycin resistance as such. This work expands the number of reliable dominant selection methods to seven, providing flexibility for the introduction of sequential genetic modifications into single strains.

Cryptococcus neoformans

Benchmark of biomarker identification and prognostic modeling methods on diverse censored data.

The practices of identifying biomarkers and developing prognostic models using genomic data has become increasingly prevalent. Such data often features characteristics that make these practices difficult, namely high dimensionality, correlations between predictors, and sparsity. Many modern methods have been developed to address these problematic characteristics while performing feature selection and prognostic modeling, but a large-scale comparison of their performances in these tasks on diverse right-censored time to event data (aka survival time data) is much needed. We have compiled many existing methods, including some machine learning methods, several which have performed well in previous benchmarks, primarily for comparison in regards to variable selection capability, and secondarily for survival time prediction on many synthetic datasets with varying levels of sparsity, correlation between predictors, and signal strength of informative predictors. For illustration, we have also performed multiple analyses on a publicly available and widely used cancer cohort from The Cancer Genome Atlas using these methods. We evaluated the methods through extensive simulation studies in terms of the false discovery rate, F1-score, concordance index, Brier score, root mean square error, and computation time. Of the methods compared, CoxBoost and the Adaptive LASSO performed well in all metrics, and the LASSO and elastic net excelled when evaluating concordance index and F1-score. The Benjamini-Hoschberg and q-value procedures showed volatile performances in controlling the false discovery rate. Some methods' performances were greatly affected by differences in the data characteristics. With our extensive numerical study, we have identified the best performing methods for a plethora of data characteristics using informative metrics. This will help cancer researchers in choosing the best approach for their needs when working with genomic data.

Humans

Targeting the F17-A Fimbrial gene: An efficient method for the quantitative detection of Escherichia coli F17.

Escherichia coli (E. coli) F17 is one of the leading bacterial causes of diarrhea in farm livestock, which cause huge economic losses and could also pose potential risks to public health. Generally, the monitoring the E. coli F17 is based on the polymerase chain reaction (PCR) and bacteria plate counting method, which were largely limited by the time-consuming nature and susceptibility to detection errors. Hence, there is an urgent need to develop a rapid and quantitative detection method for E. coli F17. In the present study, an E. coli F17 challenge experiment in ovine intestinal epithelial cells (IECs) was employed as an in vitro model. At different post-challenge time points (1&#xa0;h, 2&#xa0;h, and 3&#xa0;h), two conventional methods (bacteria plate counting and microplate method) were conducted as benchmarks to estimate the number of E. coli F17 adhering to the IECs. Additionally, total genomic DNA was extracted and quantitative Real-time PCR (qPCR) was performed to detect the relative abundance of E. coli F17 fimbrial pilin (F17-A) and adhesion (F17-G) genes. Subsequently, statistical analyses, including Pearson's correlation coefficient (PCC) method and linear curve-fitting, were performed to evaluate the correlation between the abundance of F17-A/G genes and the results of the benchmark methods. The results showed that the relative abundances of both genes were highly correlated with the number of E. coli F17 that adhered to the IECs, among them, the F17-A gene showed a stronger correlation with the bacterial counts, exhibiting a correlation coefficient&#xa0;>&#xa0;0.85. Furthermore, standard curves analyses further confirmed the out-performed quantitative performance of F17-A gene and a significantly stronger correlation with bacterial counts which exhibited an outstanding linear correlation (r&#xa0;=&#xa0;-0.9534, R2&#xa0;=&#xa0;0.9252) with amplification efficiency of 101.4%, The results of the present study indicate that targeting fimbrial genetic hallmarks via qPCR is an effective and promising method for E. coli F17 quantification, which could potentially contribute to epidemiological studies and pathogen monitoring in the livestock industry.

Detection

Evaluation of swabbing methods for culture and non-culture-based recovery of multidrug-resistant organisms from environmental surfaces.

OBJECTIVES: Sponge-Sticks (SS) and ESwabs are frequently utilized for detection of multidrug-resistant organisms (MDROs) in the environment. Head-to-head comparisons of SS and ESwabs across recovery endpoints are limited. DESIGN: We compared MDRO culture and non-culture-based recovery from (1) ESwabs, (2) cellulose-containing SS (CS), and (3)&#xa0;polyurethane-containing SS (PCS). METHODS: Known quantities of each MDRO were pipetted on a stainless-steel surface and swabbed by each method. Samples were processed, cultured, and underwent colony counting. DNA was extracted from sample eluates, quantified, and underwent metagenomic next-generation sequencing (mNGS). MDROs underwent whole genome sequencing (WGS). MDRO recovery from paired patient perirectal and PCS-collected environmental samples from clinical studies was determined. SETTING: Laboratory experiment, tertiary medical center, and long-term acute care facility. RESULTS: Culture-based recovery varied across MDRO taxa, it was highest for vancomycin-resistant Enterococcus and lowest for carbapenem-resistant Pseudomonas aeruginosa (CRPA). Culture-based recovery was significantly higher for SS compared to ESwabs except for CRPA, where all methods performed poorly. Nucleic acid recovery varied across methods and MDRO taxa. Integrated WGS and mNGS analysis resulted in successful detection of antimicrobial resistance genes, construction of high-quality metagenome-assembled genomes, and detection of MDRO genomes in environmental metagenomes across methods. In paired patient and environmental samples, multidrug-resistant Pseudomonas aeruginosa (MDRP) environmental recovery was notably poor (0/123), despite detection of MDRP in patient samples (20/123). CONCLUSIONS: Our findings support the use of SS for the recovery of MDROs. Pitfalls of each method should be noted. Method selection should be driven by MDRO target and desired endpoint.

Humans

An archaic reference-free method to jointly infer Neanderthal and Denisovan introgressed segments in modern human genomes.

Admixture between populations is a common feature of human history. Admixture events introduce new genetic variation that can fuel evolution. Characterizing the significance of admixture events on the evolution of populations across various species is of great interest to evolutionary geneticists. Local Ancestry Inference (LAI) methods infer genetic ancestry of an individual at a particular chromosomal location. Certain methods specialize in detecting archaic introgression, which consists of interbreeding between modern and archaic humans like Neanderthals and Denisovans. Most current LAI methods allow the detection of a single archaic ancestry, and post-processing may distinguish between multiple waves of introgression. These methods vary in how they choose archaic or modern reference genomes for the inference. Here, we present a new HMM-based method (DAIseg), which has the advantage of simultaneously distinguishing between multiple waves of ancient and recent admixture, using only modern human reference genomes. Simulations demonstrate that DAIseg achieves higher overall performance than state-of-the-art methods. We also apply DAIseg to Papuan populations to jointly detect Denisovan and Neanderthal introgressed segments, and identify a higher number of archaic segments than previous methods. Analysis of inferred introgressed segments, shows that we can identify evidence for two Denisovan introgression events in Papuans. Overall, on top of being able to deal with both Archaic and recent admixture, DAIseg provides a more principled approach for detecting and classifying Denisovan and Neanderthal segments which will improve downstream analysis of introgressed segments to infer the impact of archaic introgression in humans.

Denisovan

Robust and highly efficient transformation method for a minimal mycoplasma cell.

UNLABELLED: Mycoplasmas have been widely investigated for their pathogenicity, as well as for genomics and synthetic biology. Conventionally, transformation of mycoplasmas was not highly efficient, and due to the low transformation efficiency, large amounts of DNA and recipient cells were required for that purpose. Here, we report a robust and highly efficient transformation method for the minimal cell JCVI-syn3B, which was created through streamlining the genome of Mycoplasma mycoides. When the growth states of JCVI-syn3B were examined in detail by focusing on such factors as pH, color, absorbance, colony forming unit, and transformation efficiency, it was found that the growth phase after the lag phase can be divided into three distinct phases, of which the highest transformation efficiency was observed during the early exponential growth phase. Notably, the transformation efficiency of up to 4.4 &#xd7; 10-2 transformants per cell per microgram of plasmid DNA was obtained. A method to obtain several hundred to several thousand transformants with less than 0.2 mL of culture with approximately 1 &#xd7; 107-108 cells and 10 ng of plasmid DNA was developed. Moreover, a transformation method using a frozen stock of transformation-ready cells was established. These procedures and information could simplify and enhance the transformation process of minimal cells, facilitating advanced genetic engineering and biological research using minimal cells. IMPORTANCE: Mycoplasmas are parasitic and pathogenic bacteria for many animals. They are also useful bacteria to understand the cellular process of life and for bioengineering because of their simple metabolism, small genomes, and cultivability. Genetic manipulation is crucial for these purposes, but transformation efficiency in mycoplasmas is typically quite low. Here, we report a highly efficient transformation method for the minimal genome mycoplasma JCVI-syn3B. Using this method, transformants can be obtained with only 10 ng of plasmid DNA, which is around one-thousandth of the amount required for traditional mycoplasma transformations. Moreover, a convenient method using frozen stocks of transformation-ready cells was established. These improved methods play a crucial role in further studies using minimal cells.

Transformation, Bacterial

Research on multi-trait genome association study method based on Shannon information entropy.

BACKGROUND: Genetic analysis of complex traits is crucial for elucidating disease mechanisms and biological inheritance processes. However, traditional Genome-wide Association Study (GWAS) for single trait often fail to capture the synergistic effects of genetic loci on multiple traits. METHODS: This study proposes a method for analyzing the association between multiple traits and gene regions based on Shannon information entropy. Innovatively, Shannon information entropy is introduced to integrate gene region information as genetic entropy, thereby constructing an Inverse Shannon Entropy-Multi-Trait Association Analysis of Gene Region genetic model (InvSE-MTAGR). Furthermore, a partial regression test is applied to the model to establish the Inverse Partial Shannon Entropy-Multi-Trait Association Analysis of Gene Region method (InvPSE-MTAGR). When performing multi-trait analysis with InvSE-MTAGR, the method achieved statistical significance by accumulating minor effects, thereby enhancing the ability to identify pleiotropic gene regions. RESULTS: The simulation results showed that the proposed multi-trait gene region association analysis method performed well in terms of both Type I error rate control and statistical power. Leveraging tomato and sorghum datasets for validation, the proposed multi-trait gene region association analysis method based on Shannon information entropy accurately pinpointed most of the gene regions harboring candidate genes. CONCLUSION: The study reveals the advantage of multi-trait method in integrating weak-effect pleiotropic signals and capturing the correlation among traits, which provides an efficient theoretical tool for dynamic analysis of complex multi-trait genetic networks and multi-target collaborative breeding of crops.

Genome-Wide Association Study

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans

NLCD: A method to discover nonlinear causal relations among genes.

Distinguishing correlation from causation is a fundamental challenge in many scientific fields, including biology, especially when interventions like randomized controlled trials are infeasible and only observational data are available. Methods based on statistical tests of conditional independence within the Mendelian Randomization framework can detect causality between two observed variables that are each associated with a third instrumental variable. However, these methods for detecting causal relationships between traits (e.g., two gene expression or clinical traits associated with a genetic variant, all observed in the same population) often assume a linear relationship, thereby hindering the discovery of causal gene networks from genomics data. We have developed NLCD, a method for NonLinear Causal Discovery from genomics data based on nonlinear regression modeling and conditional feature importance scoring. NLCD uses these techniques to extend the statistical tests in an existing linear causal discovery method called the Causal Inference Test (CIT). We benchmarked NLCD against current state-of-the-art methods: CIT, Findr, and MRPC. On simulated datasets, NLCD performs comparably to most methods in detecting linear relations (Average AUPRC (Area Under the Precision-Recall Curve) of NLCD&#x2009;=&#x2009;0.94, CIT&#x2009;=&#x2009;0.94, Findr&#x2009;=&#x2009;0.94, and MRPC&#x2009;=&#x2009;0.99), and outperforms them in detecting nonlinear (sine and sawtooth type) relations between two genes (Average AUPRC of NLCD&#x2009;=&#x2009;0.76, CIT&#x2009;=&#x2009;0.60, Findr&#x2009;=&#x2009;0.56, and MRPC&#x2009;=&#x2009;0.73). When tested on a nonlinear subset of a yeast genomic dataset to recover known causal relations involving transcription factors, NLCD and CIT performed comparable to each other and slightly better than Findr and MRPC (Average AUPRC of NLCD&#x2009;=&#x2009;0.82, CIT&#x2009;=&#x2009;0.81, Findr&#x2009;=&#x2009;0.71, and MRPC&#x2009;=&#x2009;0.54). On application to a human genomic dataset, NLCD revealed active causal gene pairs (IRF1 &#x2192; PSME1 and HLA-C &#x2192; HLA-T) in the muscle tissue, and clarified the promises and challenges in discovering causal gene networks in tissues under in vivo human settings.

Humans