Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complex trait”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Predominant influence of environmental determinants on the persistence and avidity maturation of antibody responses to vaccines in infants.

BACKGROUND: Immune responses are complex traits influenced by genetic and environmental factors. We previously reported that genetic factors control early antibody responses to vaccines in Gambian infants. For the present study, we evaluated the determinants of the memory phase of immunoglobulin G (IgG) responses. METHODS: Antibody responses to tetanus toxoid (TT), measles vaccines, and environmental antigens (total IgG levels) were measured in 210 Gambian twin pairs recruited at birth. Intrapair correlations for monozygous and dizygous pairs were compared to estimate the environmental and genetic components of variations in response. RESULTS: In contrast to antibody responses measured in infants at age 5 months, 1 month after immunization, no significant contribution of genetic factors to anti-TT antibody and total IgG levels was detected at age 12 months. Genetic factors controlled measles antibody responses in 12-month-old infants, which indicates that the increasing influence of environmental determinants on anti-TT responses was not related to the older age of the children but, rather, to the time elapsed since immunization. Environmental factors also predominantly controlled affinity maturation and the production of high-avidity antibodies to TT. CONCLUSIONS: Genetic determinants control the early phase of the vaccine antibody response in Gambian infants, whereas environmental determinants predominantly influence antibody persistence and avidity maturation.

Aging↗

Bayesian graphical models for genomewide association studies.

As the extent of human genetic variation becomes more fully characterized, the research community is faced with the challenging task of using this information to dissect the heritable components of complex traits. Genomewide association studies offer great promise in this respect, but their analysis poses formidable difficulties. In this article, we describe a computationally efficient approach to mining genotype-phenotype associations that scales to the size of the data sets currently being collected in such studies. We use discrete graphical models as a data-mining tool, searching for single- or multilocus patterns of association around a causative site. The approach is fully Bayesian, allowing us to incorporate prior knowledge on the spatial dependencies around each marker due to linkage disequilibrium, which reduces considerably the number of possible graphical structures. A Markov chain-Monte Carlo scheme is developed that yields samples from the posterior distribution of graphs conditional on the data from which probabilistic statements about the strength of any genotype-phenotype association can be made. Using data simulated under scenarios that vary in marker density, genotype relative risk of a causative allele, and mode of inheritance, we show that the proposed approach has better localization properties and leads to lower false-positive rates than do single-locus analyses. Finally, we present an application of our method to a quasi-synthetic data set in which data from the CYP2D6 region are embedded within simulated data on 100K single-nucleotide polymorphisms. Analysis is quick (<5 min), and we are able to localize the causative site to a very short interval.

Bayes Theorem↗

Testing association between candidate-gene markers and phenotype in related individuals, by use of estimating equations.

Association studies are one of the major strategies for identifying genetic factors underlying complex traits. In samples of related individuals, conventional statistical procedures are not valid for testing association, and maximum likelihood (ML) methods have to be used, but they are computationally demanding and are not necessarily robust to violations of their assumptions. Estimating equations (EE) offer an alternative to ML methods, for estimating association parameters in correlated data. We studied through simulations the behavior of EE in a large range of practical situations, including samples of nuclear families of varying sizes and mixtures of related and unrelated individuals. For a quantitative phenotype, the power of the EE test was comparable to that of a conventional ML test and close to the power expected in a sample of unrelated individuals. For a binary phenotype, the power of the EE test decreased with the degree of clustering, as did the power of the ML test. This result might be partly explained by a modeling of the correlations between responses that is less efficient than that in the quantitative case. In small samples (< 50 families), the variance of the EE association parameter tended to be underestimated, leading to an inflation of the type I error. The heterogeneity of cluster size induced a slight loss of efficiency of the EE estimator, by comparison with balanced samples. The major advantages of the EE technique are its computational simplicity and its great flexibility, easily allowing investigation of gene-gene and gene-environment interactions. It constitutes a powerful tool for testing genotype-phenotype association in related individuals.

Female↗

Human susceptibility to viral infection: the search for HIV-protective alleles among Africans by means of genome-wide studies.

Human immunodeficiency virus (HIV) infection represents a major global health problem, with HIV now recognized as the fourth leading cause of death on a worldwide basis. One approach to developing effective anti- HIV interventions is to identify and understand the molecular mechanisms by which natural genetic variations provide protection from infection or disease progression. This approach can be used to identify human gene alleles that confer resistance or increased susceptibility to HIV infection. To date, however, this approach has been underutilized in the African population and all HIV-resistance alleles that have been described have been identified by evaluating candidate genes. This limited approach is based upon a researcher's assumption that those genes that will provide the host with a benefit can be predicted, a priori, but it does not provide for a large scale systematic screen of all possible candidate genes. Nonetheless, this method has met with some success in identifying HIV-resistance genes, mostly among the white population. The lack of a comprehensive genetic approach, both in terms of the populations studied and the percentage of the genome investigated, likely explains why all of the HIV-restriction alleles identified to date fall within two gene families, and why no resistance genes have been identified among black Africans. It is likely, as with any complex trait, that most protective alleles will provide only partial HIV resistance. Thus, HIV resistance in most persons likely arises through a QTL (quantitative trait loci) mechanism meaning that protection is a polygenic trait. This feature coupled with interpopulation genetic heterogeneity makes the candidate gene mapping approach a daunting task. A comprehensive genome-wide case-control allelic association study in the African population will maximize our chances of identifying new targets for the development of new therapeutics that have the promise of benefiting all persons infected with HIV.

Black People↗

A mutation-independent therapeutic strategem for osteogenesis imperfecta.

Given the genetically heterogeneous nature of many dominantly inherited disorders, it will be imperative to design mutation-independent therapeutic strategies to circumvent such heterogeneity. Intragenic polymorphism represents a genomic resource that may be harnessed in the development of allele-specific mutation-independent therapeutics. A hammerhead ribozyme, Rzpol1a1, selectively cleaves a common single-nucleotide polymorphism (SNP) of the human COL1A1 transcript (heterozygosity frequency of 2 pq = 0.4032, from Hardy-Weinberg equilibrium). One SNP variant contains a hammerhead ribozyme cleavage site, and the other does not. Kinetic evaluation shows Rzpol1a1 to be both specific and extremely efficient in vitro. Thus, a single efficient ribozyme has been characterized that should be valuable in the development of a gene therapy suitable for up to 1 in 5 dominant-negative osteogenesis imperfecta (OI) patients, where over 150 different mutations have been identified to date. Given the increasing characterization of intragenic SNP, it is predicted that such a mutation-independent strategy, based on selective silencing of mutant alleles at SNP, may become increasingly important in future genomics-driven drug development for many heterogeneous dominant disorders and complex traits.

Alleles↗

Effects of preconditioning and temperature during germination of 73 natural accessions of Arabidopsis thaliana.

BACKGROUND AND AIMS: Germination and establishment of seeds are complex traits affected by a wide range of internal and external influences. The effects of parental temperature preconditioning and temperature during germination on germination and establishment of Arabidopsis thaliana were examined. METHODS: Seeds from parental plants grown at 14 and at 22 degrees C were screened for germination (protrusion of radicle) and establishment (greening of cotyledons) at three different temperatures (10, 18 and 26 degrees C). Seventy-three accessions from across the entire distribution range of A. thaliana were included. KEY RESULTS: Multifactorial analyses of variances revealed significant differences in the effects of genotypes, preconditioning, temperature treatment, and their interactions on duration of germination and establishment. Reaction norms showed an enormous range of plasticity among the preconditioning and different germination temperatures. Correlations of percentage total germination and establishment after 38 d with the geographical origin of accessions were only significant for 14 degrees C preconditioning but not for 22 degrees C preconditioning. Correlations with temperature and precipitation on the origin of the accessions were mainly found at the lower germination temperatures (10 and 18 degrees C) and were absent at higher germination temperatures (26 degrees C). CONCLUSIONS: Overall, the data show huge variation of germination and establishment among natural accessions of A. thaliana and might serve as a valuable source for further germination and plasticity studies.

Adaptation, Physiological↗

Rapid differentiation of experimental populations of wheat for heading time in response to local climatic conditions.

BACKGROUND AND AIMS: Dynamic management (DM) of genetic resources aims at maintaining genetic variability between different populations evolving under natural selection in contrasting environments. In 1984, this strategy was applied in a pilot experiment on wheat (Triticum aestivum). Spatio-temporal evolution of earliness and its components (partial vernalization sensitivity, daylength sensitivity and earliness per se that determines flowering time independently of environmental stimuli) was investigated in this multisite and long-term experiment. METHODS: Heading time of six populations from the tenth generation was evaluated under different vernalization and photoperiodic conditions. KEY RESULTS: Although temporal evolution during ten generations was not significant, populations of generation 10 were genetically differentiated according to a north-south latitudinal trend for two components out of three: partial vernalization sensitivity and narrow-sense earliness. CONCLUSIONS: It is concluded that local climatic conditions greatly influenced the evolution of population earliness, thus being a major factor of differentiation in the DM system. Accordingly, a substantial proportion (approximately 25 %) of genetic variance was distributed among populations, suggesting that diversity was on average conserved during evolution but was differently distributed by natural selection (and possibly drift). Earliness is a complex trait and each genetic factor is controlled by multiple homeoalleles; the next step will be to look for spatial divergence in allele frequencies.

Climate↗

Comparative mapping in farm animals.

This paper summarises the current status of comparative mapping in farm animals. For most of the major farm animal species, a wide range of genomic tools are now available to create high-resolution genetic and physical maps of the genome. For many farm animals, the use of radiation hybrid panels and sequence data from expressed sequence tag (EST) projects has accelerated the development of high-resolution comparative maps, with human--the model species for farm animals. These tools and comparative maps are being used to map and identify the genes at the loci for simple and complex traits. The development of detailed physical maps in farm animals based on radiation hybrid panels and bacterial artificial chromosome (BAC) contigs provides a direct link between the 'information-poor' maps of farm animals and the 'information-rich' genomes of human and other model organisms.

Animals↗

Designing experiments that aid in the identification of regulatory networks.

Predictive mathematical models of the interactions of a genetic network can provide insight into the mechanisms of gene regulation, the role of various genes within a network and how multiple genes interact leading to complex traits. However, identification of the parameters and interactions is currently a limiting step in the development of such models. This work reviews the state of the art for design of experiments in biological systems and demonstrates the need for improved design of experiments through the use of a model system. Appropriate design of experiments has a profound impact on the ability to identify a model and on the quality of resulting identified model. Key issues include the selection of appropriate input sequences (e.g. random, independent multivariate inputs) and the selection of the sampling frequencies. This work demonstrates that these issues are especially important in the identification of biochemical networks and that the traditional biochemical approach is incapable of truly identifying the behavior present in such networks.

Computational Biology↗

Allelic association and disease mapping.

The application of allelic association to map genes for complex traits, particularly using high-density maps of single nucleotide polymorphisms in candidate regions, is an area of very active research. Here we present some aspects of the methodology and applications to both major gene mapping, which illustrates the effectiveness of the method, and oligogenes, where methods are still in flux and for which there have been relatively few successes to date. Several important considerations emerge, including the selection of the optimal metric for measuring association and the importance of modelling the decline in association with distance given the variability in association in a candidate region. The Malecot model of association with distance is shown to have a resolution of greater than 50 kilobases but the available evidence suggests that considerably higher resolution might be achieved with dense single nucleotide polymorphism (SNP) maps.

Alleles↗

PGS-GS: a framework integrating polygenic scores and genomic selection in animal breeding.

Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.

Animals↗

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations.

MOTIVATION: Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. RESULTS: We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (&#xd7;105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. AVAILABILITY AND IMPLEMENTATION: It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Quantitative Trait Loci↗

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗

Evaluation of epistasis detection methods for quantitative phenotypes.

MOTIVATION: Epistasis, or genetic interaction, plays a crucial role in shaping complex traits and has been increasingly recognized for its widespread influence in genetic architectures. While epistasis detection has been extensively evaluated in case-control studies, its performance with quantitative phenotypes remains comparatively understudied. RESULTS: We identified and evaluated six epistasis detection methods applicable to quantitative trait analysis: EpiSNP, Matrix Epistasis, MIDESP, PLINK Epistasis, QMDR, and REMMA. Using the EpiGEN simulator, we generated synthetic datasets modeling four classes of pairwise SNP interactions-dominant, multiplicative, recessive, and XOR. We also assessed BOOST and MDR algorithms using discretized (case-control) versions of the same datasets. Performance varied notably by interaction type: REMMA achieved the highest overall detection rate (55%), particularly excelling with dominant interactions (100%). MDR excelled with multiplicative (57%) and XOR (69%) interactions. Meanwhile, EpiSNP attained the best performance for recessive interactions (67%). All methods except BOOST produced F1 scores below 0.05 for most interaction types. We further evaluated the methods using a real-world dataset. When applied to the Adolescent Brain Cognitive Development dataset to analyse the externalizing behavior phenotype, both PLINK Epistasis and PLINK BOOST identified SNPs within the DRD2 and DRD4 genes, consistent with previously reported genetic associations. Given the variability in tool performance across interaction types, no single method provides optimal detection across all scenarios. Leveraging multiple detection algorithms may therefore yield more comprehensive insights into epistatic effects in quantitative trait analyses. AVAILABILITY AND IMPLEMENTATION: All relevant code and simulated datasets can be found at github.com/staslist/Epistasis_Review repository.

Epistasis, Genetic↗

Modelling time-varying genetic effects on binary disease risk via functional Mendelian randomization.

MOTIVATION: Genome-wide association studies have identified thousands of genetic variants associated with complex traits, establishing Mendelian randomization (MR) as a powerful framework for causal inference using variants as natural experiments. However, existing MR methods treat causal effects as static, relying on cross-sectional exposure measurements and ignoring how genetic predispositions to disease operate dynamically across the life course. Recovering age-specific causal effect functions from longitudinal data requires combining functional data representations of exposure trajectories with instrumental variable estimation strategies suitable for binary disease endpoints, a methodological gap that has remained unaddressed. RESULTS: We develop a functional MR framework for binary outcomes that integrates functional principal component analysis with two-stage residual inclusion (2SRI), ensuring consistent estimation under the nonlinear logistic link function that renders standard instrumental variable estimators inconsistent. Simulations across different causal effect trajectory shapes, varying measurement densities, and varying instrument strengths demonstrate accurate recovery of time-varying genetically predicted effects with minimal bias. Applied to UK Biobank data, the framework identifies an age-specific causal effect of genetically predicted body mass index on type 2 diabetes risk concentrated in early mid-adulthood and progressively attenuating thereafter. Concordance between the proposed 2SRI estimator applied to type 2 diabetes and the established continuous-outcome functional MR estimator applied to the paired glycated haemoglobin marker in the same cohort provides indirect empirical support for the validity of the proposed approach. AVAILABILITY AND IMPLEMENTATION: The method is implemented in the R package mvfmr, with a full tutorial vignette.

Mendelian Randomization Analysis↗

A module-based approach for post-omics, post-GWAS network-based gene classification.

MOTIVATION: Complex traits and diseases are highly polygenic and understanding the full set of genes involved is a central challenge in biomedicine. However, due to sample size limitations and noise (technical and biological), experimental approaches for disease-gene discovery such as transcriptomics and GWAS result in long, noisy, heterogeneous gene lists, which may be trimmed to a subset of likely relevant genes while leaving several false negatives. Computational gene classification approaches, especially those using genome-scale molecular interaction networks, are promising avenues for complementing such experimental findings by analytically expanding observed gene lists based on the functional relatedness between genes. We previously introduced the network-based gene classification approach, GenePlexus, which was rigorously benchmarked to show state-of-the-art performance, especially for predicting novel genes associated with biological processes and fine-grained phenotypes. Network-based gene classification performance,however, declines for diseases, especially when the inputs are omics and GWAS-based long gene lists. RESULTS: Here, we show that these disease gene lists span multiple biological processes spread across the molecular network, and we propose ModGenePlexus, a new network-based gene classification method that takes a two-stage approach. First, clustering and semi-supervised learning decomposes the input gene list into coherent, denoised network gene modules. Then, ModGenePlexus trains supervised (GenePlexus) classifiers for each module and aggregates predictions to return genome-wide rankings. We benchmarked ModGenePlexus across simulated data, transcriptomic signatures, and GWAS datasets (together spanning hundreds of diseases), showing improved recovery of known disease genes compared to GenePlexus. Beyond improved classification, the results of enrichment analysis of ModGenePlexus outputs are much more interpretable by virtue of revealing nuanced biological processes. Together, these results establish ModGenePlexus as a scalable, interpretable tool for gene classification of GWAS- and omics-derived gene lists across diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: ModGenePlexus is freely available on GitHub at https://github.com/krishnanlab/ModGenePlexus, and the full source code and results supporting this study are available on Zenodo at https://zenodo.org/records/19857910.

Genome-Wide Association Study↗

SNPannotator: automated functional annotation of genetic variants and linked proxies.

SUMMARY: Genome-wide association studies (GWASs) have identified thousands of genetic variants associated with complex traits and diseases. However, explaining the mechanisms underlying phenotypic variation remains challenging. Here, we introduce SNPannotator, an automated post-GWAS analysis software package designed to streamline the interpretation of GWAS findings. Our pipeline implements a multi-step process that identifies proxy variants in high linkage disequilibrium (LD) with associated lead variants, then queries comprehensive resources (including Ensembl, the GTEx Portal, the eQTL Catalog, and STRING DB) for genomic position, deleteriousness, regulatory annotations, clinical significance, trait associations, expression (eQTLs) and splicing quantitative trait loci (sQTLs), and functional enrichment analyses and compiles the results into user-friendly reports. This package is implemented in the R programming language and includes auxiliary functions for variant lookup and LD exploration. SNPannotator provides a practical framework for efficiently deriving biologically meaningful insights from GWAS data and for assisting researchers in prioritizing candidate variants for functional validation. AVAILABILITY AND IMPLEMENTATION: The SNPannotator package is available from the Comprehensive R Archive Network (CRAN) at https://cran.r-project.org/web/packages/SNPannotator. The development version and tutorial is available on GitHub (https://github.com/omicslaboratory/SNPannotator). The online version of the package is available at https://omicslab.org/snpannotator.

Software↗

Association mapping and fine mapping with TreeLD.

SUMMARY: The program package TreeLD implements a unified approach to association mapping and fine mapping of complex trait loci and a novel approach to visualizing association data, based on an inferred ancestry of the sample. Fundamentally, the TreeLD approach is based on the idea that the evidence for association at a particular position is contained in the ancestral tree relating the sampled chromosomes at that position. TreeLD provides an easy-to-use interface and can be applied to case-control, TDT trio and quantitative trait data.

Algorithms↗