Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “complex trait”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Gene-environment interaction analysis in atopic eczema: evidence from large population datasets and modelling in vitro.

BACKGROUND: Environmental factors play a role in the pathogenesis of complex traits including atopic eczema (AE) and a greater understanding of gene-environment interactions (G*E) is needed to define pathomechanisms for disease prevention. We analysed data from 16 European studies to test for interaction between the 24 most significant AE-associated loci identified from genome-wide association studies and 18 early-life environmental factors. We tested for replication using a further 10 studies and in vitro modelling to independently assess findings. RESULTS: The discovery analysis showed suggestive evidence for interaction (p<0.05) between 7 environmental factors (antibiotic use, cat ownership, dog ownership, breastfeeding, elder sibling, smoking and washing practices) and at least one established variant for AE, 14 interactions in total (maxN=25,339). In replication analysis (maxN=252,040) dog exposure*rs10214237 (on chromosome 5p13.2 near IL7R) was nominally significant (ORinteraction=0.91 [0.83-0.99] P=0.025), with a risk effect of the T allele observed only in those not exposed to dogs. A similar interaction with rs10214237 was observed for siblings in the discovery analysis (ORinteraction=0.84[0.75-0.94] P=0.003), but replication analysis was under-powered ORinteraction=1.09[0.82-1.46]). Rs10214237 homozygous risk genotype is associated with lower IL-7R expression in human keratinocytes, and dog exposure modelled in vitro showed a differential response according to rs10214237 genotype. CONCLUSIONS: Interaction analysis and functional assessment provide evidence that early-life dog exposure may modify the genetic effect of rs10214237 on AE via IL7R, supporting observational epidemiology showing a protective effect for dog ownership. The lack of evidence for other G*E studied here implies that only weak effects are likely to occur.

Atopic eczema↗

Genome reorganisation and expansion shape 3D genome architecture and define a distinct regulatory landscape in coleoid cephalopods.

How genomic changes translate into organismal novelties is often confounded by the multi-layered nature of genome architecture and the long evolutionary timescales over which molecular changes accumulate. Coleoid cephalopods (squid, cuttlefish, and octopus) provide a unique system to study these processes due to a large-scale chromosomal rearrangement in the coleoid ancestor that resulted in highly modified karyotypes, followed by lineage-specific fusions, translocations, and repeat expansions. How these events have shaped gene regulatory patterns underlying the evolution of coleoid innovations, including their large and elaborately structured nervous systems, novel organs, and complex behaviours, remains poorly understood. To address this, we integrate Micro-C, RNA-seq, and ATAC-seq across multiple coleoid species, developmental stages, and tissues. We find that while topological compartments are broadly conserved, hundreds of chromatin loops are species- and context-specific, with distinct regulation signatures and dynamic expression profiles. CRISPR-Cas9 knockout of a putative regulatory sequence within a conserved region demonstrates the role of loops in neural development and the prevalence of long-range, inter-compartmental interactions. We propose that differential evolutionary constraints across the coleoid 3D genome allow macroevolutionary processes to shape genome topology in distinct ways, facilitating the emergence of novel regulatory entanglements and ultimately contributing to the evolution and maintenance of complex traits in coleoids.

Journal Article↗

Predicting emergent phenotypes from single cell populations using CELLECTION.

Biological systems exhibit emergent phenotypes that arise from the collective behavior of individual components, such as whole-organ functions that arise from the coordinated activity of its individual cells, or organism-level phenotypes that result from the functional interplay of collections of genes in the genome. We present CELLECTION, a deep learning framework that learns to associate subgroups of instances with different emergent phenotypes. We show CELLECTION enables interpretable predictions for heterogeneous tasks, including disease classification, identification of disease-associated cell subtypes, alignment of developmental stages between human model systems, and even predicting relative hand-wing indices across the avian lineage. CELLECTION therefore provides a scalable and flexible framework for identifying key cellular or genetic signatures underlying complex traits in development, disease, and evolution.

Journal Article↗

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗

Extended intermarker linkage disequilibrium in the Afrikaners.

In this study we conducted an investigation of the background level of linkage disequilibrium (LD) in the Afrikaner population to evaluate the appropriateness of this genetic isolate for mapping complex traits. We analyzed intermarker LD in 62 nuclear families using microsatellite markers covering extended chromosomal regions. The markers were selected to allow the first direct comparison of long-range LD in the Afrikaners to LD in other demographic groups. Using several statistical measures, we find significant evidence for LD in the Afrikaners extending remarkably over a 6-cM range. In contrast, LD decays significantly beyond 3-cM distances in the other founder and outbred populations examined. This study strongly supports the appropriateness of the Afrikaner population for genome-wide scans that exploit LD to map common, multigenic disorders.

Chromosome Mapping↗

Genetic analysis of case/control data using estimated haplotype frequencies: application to APOE locus variation and Alzheimer's disease.

There is growing debate over the utility of multiple locus association analyses in the identification of genomic regions harboring sequence variants that influence common complex traits such as hypertension and diabetes. Much of this debate concerns the manner in which one can use the genotypic information from individuals gathered in simple sampling frameworks, such as the case/control designs, to actually assess the association between alleles in a particular genomic region and a trait. In this paper we describe methods for testing associations between estimated haplotype frequencies derived from multilocus genotype data and disease endpoints assuming a simple case/control sampling design. These proposed methods overcome the lack of phase information usually associated with samples of unrelated individuals and provide a comprehensive way of assessing the relationship between sequence or multiple-site variation and traits and diseases within populations. We applied the proposed methods in a study of the relationship between polymorphisms within the APOE gene region and Alzheimer's disease. Cases and controls for this study were collected from the United States and France. Our results confirm the known association between the APOE locus and Alzheimer's disease, even when the epsilon 4 polymorphism is not contained in the tested haplotypes. This suggests that, in certain situations, haplotype information and linkage disequilibrium-induced associations between polymorphic loci that neighbor loci harboring functional sequence variants can be exploited to identify disease-predisposing alleles in large, freely mixing populations via estimated haplotype frequency methods.

Aged↗

Sequence diversity in genes of lipid metabolism.

Elevated plasma lipoprotein levels play a crucial role in the development of coronary artery disease. Genetic factors strongly influence the levels of plasma lipoproteins, but the genes and sequence variations contributing to the most common forms of dyslipidemias are not known. We used GeneChip probe arrays to resequence the coding regions of 10 key genes of lipid metabolism. The sequences of these genes were analyzed in 80 dyslipidemic individuals. Fourteen nonsynonymous and twenty-two synonymous single nucleotide changes were identified that could be confirmed by conventional sequencing. Seven of the fourteen nonsynonymous sequence variants were polymorphisms with allele frequency >1% in the general population. The remaining seven were not found in normolipidemic controls (25 Caucasians and 25 African-Americans). The relationship between nonsynonymous sequence variations and various dyslipidemias was explored in association and family studies. No evidence was found for coding sequence variations in any of the 10 genes contributing to dyslipidemia. Only a single sequence variation, a missense mutation in the low density lipoprotein receptor gene, co-segregated with hyperlipidemia in the proband's family. This study illustrates some of the difficulties associated with identifying sequence variations contributing to complex traits.

Adult↗

Linkage disequilibrium between microsatellite markers extends beyond 1 cM on chromosome 20 in Finns.

Linkage disequilibrium (LD) is a proven tool for evaluating population structure and localizing genes for monogenic disorders. LD-based methods may also help localize genes for complex traits. We evaluated marker-marker LD using 43 microsatellite markers spanning chromosome 20 with an average density of 2.3 cM. We studied 837 individuals affected with type 2 diabetes and 386 mostly unaffected spouse controls. A test of homogeneity between the affected individuals and their spouses showed no difference, allowing the 1223 individuals to be analyzed together. Significant (P < 0.01) LD was observed using a likelihood ratio test in all (11/11) marker pairs within 1 cM, 78% (25/32) of pairs 1-3 cM apart, and 39% (7/18) of pairs 3-4 cM apart, but for only 12 of 842 pairs more than 4 cM apart. We used the human genome project working draft sequence to estimate kilobase (kb) intermarker distances, and observed highly significant LD (P < 10(-10)) for all six marker pairs up to 350 kb apart, although the correlation of LD with cM is slightly better than the correlation with megabases. These data suggest that microsatellites present at 1-cM density are sufficient to observe marker-marker LD in the Finnish population.

Alleles↗

High-throughput variation detection and genotyping using microarrays.

The genetic dissection of complex traits may ultimately require a large number of SNPs to be genotyped in multiple individuals who exhibit phenotypic variation in a trait of interest. Microarray technology can enable rapid genotyping of variation specific to study samples. To facilitate their use, we have developed an automated statistical method (ABACUS) to analyze microarray hybridization data and applied this method to Affymetrix Variation Detection Arrays (VDAs). ABACUS provides a quality score to individual genotypes, allowing investigators to focus their attention on sites that give accurate information. We have applied ABACUS to an experiment encompassing 32 autosomal and eight X-linked genomic regions, each consisting of approximately 50 kb of unique sequence spanning a 100-kb region, in 40 humans. At sufficiently high-quality scores, we are able to read approximately 80% of all sites. To assess the accuracy of SNP detection, 108 of 108 SNPs have been experimentally confirmed; an additional 371 SNPs have been confirmed electronically. To access the accuracy of diploid genotypes at segregating autosomal sites, we confirmed 1515 of 1515 homozygous calls, and 420 of 423 (99.29%) heterozygotes. In replicate experiments, consisting of independent amplification of identical samples followed by hybridization to distinct microarrays of the same design, genotyping is highly repeatable. In an autosomal replicate experiment, 813,295 of 813,295 genotypes are called identically (including 351 heterozygotes); at an X-linked locus in males (haploid), 841,236 of 841,236 sites are called identically.

Algorithms↗

Trimming, weighting, and grouping SNPs in human case-control association studies.

The search for genes underlying complex traits has been difficult and often disappointing. The main reason for these difficulties is that several genes, each with rather small effect, might be interacting to produce the trait. Therefore, we must search the whole genome for a good chance to find these genes. Doing this with tens of thousands of SNP markers, however, greatly increases the overall probability of false-positive results, and current methods limiting such error probabilities to acceptable levels tend to reduce the power of detecting weak genes. Investigating large numbers of SNPs inevitably introduces errors (e.g., in genotyping), which will distort analysis results. Here we propose a simple strategy that circumvents many of these problems. We develop a set-association method to blend relevant sources of information such as allelic association and Hardy-Weinberg disequilibrium. Information is combined over multiple markers and genes in the genome, quality control is improved by trimming, and an appropriate testing strategy limits the overall false-positive rate. In contrast to other available methods, our method to detect association to sets of SNP markers in different genes in a real data application has shown remarkable success.

Case-Control Studies↗

Linkage disequilibrium and haplotype diversity in the genes of the renin-angiotensin system: findings from the family blood pressure program.

Association studies of candidate genes with complex traits have generally used one or a few single nucleotide polymorphisms (SNPs), although variation in the extent of linkage disequilibrium (LD) within genes markedly influences the sensitivity and precision of association studies. The extent of LD and the underlying haplotype structure for most candidate genes are still unavailable. We sampled 193 blacks (African-Americans) and 160 whites (European-Americans) and estimated the intragenic LD and the haplotype structure in four genes of the renin-angiotensin system. We genotyped 25 SNPs, with all but one of the pairs spaced between 1 and 20 kb, thus providing resolution at small scale. The pattern of LD within a gene was very heterogeneous. Using a robust method to define haplotype blocks, blocks of limited haplotype diversity were identified at each locus; between these blocks, LD was lost owing to the history of recombination events. As anticipated, there was less LD among blacks, the number of haplotypes was substantially larger, and shorter haplotype segments were found, compared with whites. These findings have implications for candidate-gene association studies and indicate that variation between populations of European and African origin in haplotype diversity is characteristic of most genes.

Adult↗

The linkage disequilibrium maps of three human chromosomes across four populations reflect their demographic history and a common underlying recombination pattern.

The extent and patterns of linkage disequilibrium (LD) determine the feasibility of association studies to map genes that underlie complex traits. Here we present a comparison of the patterns of LD across four major human populations (African-American, Caucasian, Chinese, and Japanese) with a high-resolution single-nucleotide polymorphism (SNP) map covering almost the entire length of chromosomes 6, 21, and 22. We constructed metric LD maps formulated such that the units measure the extent of useful LD for association mapping. LD reaches almost twice as far in chromosome 6 as in chromosomes 21 or 22, in agreement with their differences in recombination rates. By all measures used, out-of-Africa populations showed over a third more LD than African-Americans, highlighting the role of the population's demography in shaping the patterns of LD. Despite those differences, the long-range contour of the LD maps is remarkably similar across the four populations, presumably reflecting common localization of recombination hot spots. Our results have practical implications for the rational design and selection of SNPs for disease association studies.

Black or African American↗

Resistance to salmonellosis in the chicken is linked to NRAMP1 and TNC.

Natural resistance to infection with Salmonella typhimurium in mice is controlled by two major loci, Bcg and Lps, located on mouse chromosomes 1 and 4, respectively. Both Bcg and Lps exert pleiotropic effects and contribute to cytostatic/cytocidal activities of the macrophage. Bcg encodes for a membrane phosphoglycoprotein designated Nrampl (natural resistance-associated macrophage protein 1), which belongs to an ancient family of membrane proteins, Lps has not been cloned yet, but its location on mouse chromosome 4 has been refined for positional cloning. As in mice, chicken inbred lines differ in their susceptibility to infection with Salmonella typhimurium. We have tested the candidacy of the chicken homologs of Nrampl and Tnc (a locus closely linked to Lps), in the differential resistance of chicken inbred lines to infection with S. typhimurium. We have first analyzed six inbred chicken lines of Salmonella-resistant or Salmonella-susceptible phenotypes for the presence of nucleotide sequence variations within the coding portion of NRAMP1. We have identified 11 sequence variations within NRAMP1 in the chicken inbred lines tested: 10 of these represented either silent mutations or conservative changes. However, one G-->A substitution at nucleotide 696 resulted in the nonconservative replacement of Arg223 to Gln223 within the predicted TM5-6 region. This allelic variant was specific to the susceptible line C and not observed in any of the resistant strains. To investigate the effect of NRAMP1 and TNC on resistance to infection with S. typhimurium, 425 (W1 x C)F1 x C chicken progeny were examined during a period of 15 days postinfection. Together, NRAMP1 and TNC explain 33% of the early differential resistance to infection with S. typhimurium of parental lines C and W1. Our data established that resistance to infection with S. typhimurium in chickens is inherited as a complex trait and that comparative mapping has proven to be useful to identify Salmonella-resistance genes in the chicken.

Amino Acid Sequence↗

A homogeneous, ligase-mediated DNA diagnostic test.

Single-nucleotide variations are the most widely distributed genetic markers in the human genome. A subset of these variations, the substitution mutations, are responsible for most genetic disorders. As single nucleotide polymorphism (SNP) markers are being developed for molecular diagnosis of genetic disorders and large-scale population studies for genetic analysis of complex traits, a simple, sensitive, and specific test for single nucleotide changes is highly desirable. In this report we describe the development of a homogeneous DNA detection method that requires no further manipulations after the initial reaction is set up. This assay, named dye-labeled oligonucleotide ligation (DOL), combines the PCR and the oligonucleotide ligation reaction in a two-stage thermal cycling sequence with fluorescence resonance energy transfer (FRET) detection monitored in real time. Because FRET occurs only when the donor and acceptor dyes are in close proximity, one can infer the genotype or mutational status of a DNA sample by monitoring the specific ligation of dye-labeled oligonucleotide probes. We have successfully applied the DOL assay to genotype 10 SNPs or mutations. By designing the PCR primers and ligation probes in a consistent manner, multiple assays can be done under the same thermal cycling conditions. The standardized design and execution of the DOL assay means that it can be automated for high-throughput genotyping in large-scale population studies.

DNA Ligases↗

Microsatellite marker content mapping of 12 candidate genes for obesity: assembly of seven obesity screening panels for automated genotyping.

Twin studies, adoption studies, and studies of familial aggregation indicate that obesity has a genetic component. Whereas, the genetic factors predisposing to obesity have been elucidated for several rare syndromes, the factors responsible for obesity in the general population have remained elusive. Genetic studies of complex traits are often accelerated by the use of candidate genes. To facilitate genetic studies of human obesity, seven multiplex panels of candidate genes for obesity that are suitable for fluorescent genotyping have been assembled. The multiplex panels are composed of 66 microsatellite markers linked tightly to 16 human gene products that are of potential importance in the control of body weight or linked to syndromic forms of obesity. As part of these efforts 12 previously cloned genes have been placed on the human physical map. In addition the chromosomal location of three of these genes, ART, NYP Y6R, and PPARgamma, are reported for the first time. These resources will be of use in studies to identify the genetic factors responsible for human obesity. [Figures are available at http://www.genome.org]

Agouti-Related Protein↗

The physiology and biophysics of an aluminum tolerance mechanism based on root citrate exudation in maize.

Al-induced release of Al-chelating ligands (primarily organic acids) into the rhizosphere from the root apex has been identified as a major Al tolerance mechanism in a number of plant species. In the present study, we conducted physiological investigations to study the spatial and temporal characteristics of Al-activated root organic acid exudation, as well as changes in root organic acid content and Al accumulation, in an Al-tolerant maize (Zea mays) single cross (SLP 181/71 x Cateto Colombia 96/71). These investigations were integrated with biophysical studies using the patch-clamp technique to examine Al-activated anion channel activity in protoplasts isolated from different regions of the maize root. Exposure to Al nearly instantaneously activated a concentration-dependent citrate release, which saturated at rates close to 0.5 nmol citrate h(-1) root(-1), with the half-maximal rates of citrate release occurring at about 20 microM Al(3+) activity. Comparison of citrate exudation rates between decapped and capped roots indicated the root cap does not play a major role in perceiving the Al signal or in the exudation process. Spatial analysis indicated that the predominant citrate exudation is not confined to the root apex, but could be found as far as 5 cm beyond the root cap, involving cortex and stelar cells. Patch clamp recordings obtained in whole-cell and outside-out patches confirmed the presence of an Al-inducible plasma membrane anion channel in protoplasts isolated from stelar or cortical tissues. The unitary conductance of this channel was 23 to 55 pS. Our results suggest that this transporter mediates the Al-induced citrate release observed in the intact tissue. In addition to the rapid Al activation of citrate release, a slower, Al-inducible increase in root citrate content was also observed. These findings led us to speculate that in addition to the Al exclusion mechanism based on root citrate exudation, a second internal Al tolerance mechanism may be operating based on Al-inducible changes in organic acid synthesis and compartmentation. We discuss our findings in terms of recent genetic studies of Al tolerance in maize, which suggest that Al tolerance in maize is a complex trait.

Adaptation, Physiological↗

Dissection of maize kernel composition and starch production by candidate gene association.

Cereal starch production forms the basis of subsistence for much of the world's human and domesticated animal populations. Starch concentration and composition in the maize (Zea mays ssp mays) kernel are complex traits controlled by many genes. In this study, an association approach was used to evaluate six maize candidate genes involved in kernel starch biosynthesis: amylose extender1 (ae1), brittle endosperm2 (bt2), shrunken1 (sh1), sh2, sugary1, and waxy1. Major kernel composition traits, such as protein, oil, and starch concentration, were assessed as well as important starch composition quality traits, including pasting properties and amylose levels. Overall, bt2, sh1, and sh2 showed significant associations for kernel composition traits, whereas ae1 and sh2 showed significant associations for starch pasting properties. ae1 and sh1 both associated with amylose levels. Additionally, haplotype analysis of sh2 suggested this gene is involved in starch viscosity properties and amylose content. Despite starch concentration being only moderately heritable for this particular panel of diverse maize inbreds, high resolution was achieved when evaluating these starch candidate genes, and diverse alleles for breeding and further molecular analysis were identified.

Base Sequence↗