Search PubMedSearch

SEARCH · Search PubMed

Results for “Genotype Data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Integrating plant phenotypic and genotypic data in the AGENT project: a BrAPI service implementation.

MOTIVATION: The AGENT project established a network of actively cooperating European genebanks, integrating genomic and phenotypic data from accessions of wheat and barley. Due to specific storage demands for phenotypic and genotypic data, the project used separate database instances and backend technologies to manage integrated phenotypic and genotypic data. RESULTS: We discuss the challenges encountered when integrating dispersed data to serve through a single interface such as the Plant Breeding Application Programming Interface, BrAPI. We examine how the consistent mappability of genebank data to the BrAPI model can enable the implementation of effective services. The advantages of BrAPI in transparently linking distributed data entities through embedded, unique identifiers are highlighted. We present a technical solution involving a BrAPI proxy, which combines and merges separate BrAPI endpoints. Finally, we demonstrate the AGENT BrAPI implementation with an illustrative example that validates a suggested SNP for a trait from the literature by linking phenotypic, genotypic and passport data. AVAILABILITY AND IMPLEMENTATION: The BrAPI proxy implementation and documentation is available at the Python Package Index (https://pypi.org/project/brapi-proxy) and archived in Zenodo (doi: 10.5281/zenodo.19436445). SUPPLEMENTARY INFORMATION: A Jupyter Notebook file for the validation example using a marker-trait relationship found in the literature.

Phenotype

Inferences about linkage disequilibrium.

Existing theory for inferences about linkage disequilibrium is restricted to a measure defined on gametic frequencies. Unless gametic frequencies are directly observable, they are inferred from genotypic frequencies under the assumption of random union of gametes. Primary emphasis in this paper is given to genotypic data, and disequilibrium coefficients are defined for all subsets of two or more of the four genes, two at each of two loci, carried by an individual. Linkage disequilibrium coefficients are defined for genes within and between gametes, and methods of estimating and testing these coefficients are given for gametic data. For genotypic data, when coupling and repulsion double heterozygotes cannot be distinguished. Burrows' composite measure of linkage disequilibrium is discussed. In particular, the estimate for this measure and hypothesis tests based on it are compared to the usual maximum likelihood estimate of gametic linkage disequilibrium, and corresponding likelihood ratio or contingency chi-square tests. General use of the composite measure, whether or not random union of gametes is an appropriate assumption, is recommended. Attention is given to small samples, where the non-normality of gene frequencies will have greatest effect on methods of inference based on normal theory. Even tools such as Fisher's z-transformation for the correlation of gene frequencies are found to perform quite satisfactorily.

Gene Frequency

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+ approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+ led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+ also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+ approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+ can be integrated into existing R pipelines for commercial methylation arrays, such as 450 K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Clinically Relevant Pharmacogenomic Variant Frequencies in Kazakh, Russian, and Uzbek Population Groups Residing in Kazakhstan.

Central Asian populations remain underrepresented in pharmacogenomic research, limiting the availability of population-specific data for genotype-informed prescribing and precision medicine. This study analyzed clinically relevant pharmacogenomic variant frequencies in Kazakh, Russian, and Uzbek population groups residing in Kazakhstan using genome-wide genotype data from 1301 individuals: Kazakh (n = 1111), Russian (n = 156), and Uzbek (n = 34). ClinPGx, a PharmGKB-based clinical annotation framework that prioritizes variant-drug associations according to levels of evidence, was used to select variants with evidence levels 1A, 1B, and 2A. In total, 112 directly genotyped variants were retained for population-specific allele and genotype frequency analysis. All 112 variants were queried against the gnomAD v4.1 genome and exome reference datasets. Of these, matching allele-frequency data for the predefined reported allele were available in at least one of the two gnomAD datasets for 103 variants, whereas for 9 variants the VEP-based query did not return a matching gnomAD frequency for that allele. Frequencies were reported for the same predefined reported allele across all groups, and differences between the study groups were assessed using 95% confidence intervals, Fisher's exact tests, and false discovery rate correction. Genotype counts and the proportions of individuals carrying at least one copy of the reported allele were also summarized for all selected variants. Several pharmacogenomic variants showed population-specific frequency patterns, including NUDT15 rs116855232, SLCO1B1 rs4149056, VKORC1 rs9934438, and UGT1A1 rs10929302. Comparison with gnomAD showed that the observed frequencies were variant-specific and could not be consistently approximated by a single broad genetic ancestry group. Reference-based population structure analysis provided additional ancestry context and supported separate reporting by population group. The study did not evaluate clinical outcomes or make individual prescribing recommendations, and the small Uzbek sample size limits the precision of frequency estimates for this group, particularly for rare variants. Overall, this study provides a clinically prioritized pharmacogenomic frequency resource for underrepresented population groups in Kazakhstan and supports broader Central Asian representation in pharmacogenomic implementation research.

Central Asia

A haplotype-based 'haplotype relative risk' approach to detecting allelic associations.

A novel variation of the Haplotype Relative Risk (HRR) of Rubinstein et al. [Hum Immunol 1981;3:384] is proposed, in order to glean increased information about linkage disequilibrium or allelic associations by analyzing haplotype-based data rather than genotypic data. It is shown that statistical tests based on our design give much higher power than those based on the original HRR approach. Several additional nonparametric tests based on the same data are analyzed, and power is computed for each of them. Further, parametric likelihood methods are applied to testing linkage equilibrium, and estimating delta, the coefficient of linkage disequilibrium, from the same data.

Genetic Linkage

CRISPR-Cas9-induced genetic mosaicism in three species of the microcrustacean Daphnia.

Genetic mosaicism can arise from in vivo CRISPR-Cas9 gene editing, especially in the embryos. This study evaluates the extent of genetic mosaicism resulted from CRISPR-Cas9-mediated knockout for 11 genes in the freshwater microcrustacean Daphnia magna, Daphnia pulex, and Daphnia sinensis. Based on extensive genotyping data of the asexually produced progenies of successfully edited females, we find strong evidence of mosaicism in 9 of these genes. The genotyping data also suggest that the gene editing activity can take place as early as the one-cell embryo stage and extends into the 32-cell and later stages. This study establishes genetic mosaicism as an important feature of Cas9-mediated gene editing in Daphnia.

Animals

Population genetics of hypervariable loci: analysis of PCR based VNTR polymorphism within a population.

Using a polymerase chain reaction (PCR) based method, genotypes at two hypervariable loci (3' to the Apo-B-structural gene and at the ApoC-II gene) were determined by size classification of alleles. Genotype data at the Apo-B locus (Apo-B VNTR) were obtained on 240 French Caucasians; the sample size for the ApoC-II VNTR was 162. For 160 individuals two-locus genotype data were available. Applications of some recently developed statistical methods to these data indicate that both of these loci are at Hardy-Weinberg equilibrium (HWE) and there is no indication of allelic associations between these two unlinked loci. In addition, the observed numbers of alleles (12 for the Apo-B and 11 for the ApoC-II VNTR loci) are also consistent with their respective expectations based on the observed heterozygosities (76.9% for the Apo-B and 85.9% for the ApoC-II loci) suggesting genetic homogeneity of this population-based sample. The multimodal distribution of allele sizes observed for both loci indicate that the production of new alleles at such VNTR loci may be caused by more than one molecular mechanism. The utility of such highly polymorphic loci for human genetic research and forensic applications are discussed in the context of these findings.

Alleles

Linkage investigation of three putative tuberous sclerosis determining loci on chromosomes 9q, 11q, and 12q. The Tuberous Sclerosis Collaborative Group.

Previous linkage studies in tuberous sclerosis have implicated three disease determining loci at 9q, 11q, and 12q. We have collated phenotypic and genotypic data on 1622 members of 128 families with tuberous sclerosis in order to evaluate simultaneously the evidence for these putative loci. Affection status in the family members has been reassessed using uniform diagnostic criteria and genotypic data extensively checked before analysis under alternative models of locus heterogeneity. One tuberous sclerosis determining locus, accounting for approximately 50% of the families studied, has been found to map in the region of D9S10 on 9q34 but no evidence has been found to support the existence of major loci on 11q or 12q. A locus, or loci, elsewhere in the genome is likely to account for tuberous sclerosis in most non-chromosome 9 linked families.

Chromosome Mapping

Genome-wide cis-expression Quantitative Trait Loci (eQTL) and transcriptomic signals reveal distinct molecular regulation across correlated feed efficiency traits.

INTRODUCTION: Feed efficiency (FE) is a complex trait which determines livestock production profitability, yet the molecular mechanisms behind it remain unclear. This study investigated the blood transcriptomic profile of lambs, alongside genotype data with the aim to uncover the genetic basis of FE traits such as absolute dry matter intake (DMIabsolute), DMI adjusted for body size (DMIadjusted), average daily live weight gain (ADG), and residual feed intake (RFI). MATERIALS AND METHODS: Bulk RNA-Seq and genotype data were analysed using three complementary approaches: differential gene expression (DGE) analysis, weighted gene co-expression network analysis (WGCNA), and cis-expression Quantitative Trait Loci (cis-eQTL) mapping. These methods were used independently to identify genes and regulatory networks associated with FE traits and to investigate evidence supporting multi-trait candidate gene selection. RESULTS: DGE analysis revealed 2, 24, 85 and 4 differentially expressed genes for DMIabsolute, DMIadjusted, ADG, and RFI (Padjusted < 0.05), functionally enriched in sensory perception, ATP-dependent chromatin remodeling, Notch signaling and immune response pathways. 9 gene modules significantly associated with the FE traits (P &#x2264; 0.05) with correlations ranging from r = -0.56 to 0.49, were identified using WGCNA. Single nucleotide polymorphism (SNP)-level cis-eQTL analysis identified 93 eSNPs associated with 74 genes (false discovery rate (FDR) < 0.05), while permutation-derived gene level analysis identified 280 eGenes (FDR < 0.2, empirical P < 0.03). Across the three analyses, applying thresholds of DGE (Padjusted < 0.05), WGCNA (correlation, P &#x2264; 0.05), and cis-eQTL gene-level significance (empirical P < 0.05), multiple overlapping genes were identified including DNMT3A, KANSL1, NCOR1 for DMIadjusted, ACOX2, FANCF, CIMIP2B, LOC101115106, ARMH2, LOC132657496 for ADG, and LOC114114576 for RFI representing regulators of variations in FE. DISCUSSION: The integration of DGE, WGCNA, and cis-eQTL analyses identified key genes and regulatory mechanisms associated with variation in FE traits. These results highlight that integrated multi-trait candidate gene identification approaches can reveal key genes that lower feed intake while maintaining animal growth, supporting breeding strategies aimed at improving efficiency and long-term economic sustainability in sheep.

average daily gain (ADG)

Understanding recurrence in Mycobacterium avium complex pulmonary disease: genotypic strategies to support clinical decision-making.

Pulmonary disease caused by Mycobacterium avium complex (MAC-PD) is a chronic, recurrent disease, and its high recurrence rate after treatment makes clinical management difficult. Distinguishing whether recurrence is due to persistence of existing strains or reinfection with new strains is essential for establishing treatment strategies, preventing overuse of antimicrobials, and establishing infection control measures. According to reports, 54%-74% of MAC-PD recurrence is due to reinfection, which may be mainly related to environmental reservoirs such as household water supply. In this review, we present various clinical scenarios in which MAC-PD recurrence may occur and examine genotyping techniques as a strategy to distinguish and respond to them. From traditional methods such as IS1245-based restriction fragment length polymorphism, pulsed-field gel electrophoresis, and hsp65 and rpoB gene sequencing to high-resolution analysis techniques such as multilocus sequence testing and whole-genome sequencing, the latest molecular typing methods are comprehensively summarized. Integrating these genotype data into clinical settings, standardizing single-nucleotide polymorphism-based interpretation thresholds, and promoting the establishment of a global MAC strain database will make a substantial contribution to more accurately distinguishing the recurrence mechanisms of MAC-PD and establishing personalized treatment strategies.IMPORTANCEThe global burden of nontuberculous mycobacterial pulmonary disease (PD) is increasing, with Mycobacterium avium (MAC)-PD being the most prevalent and clinically challenging form. Its low treatment success rates, high frequency of recurrence, and persistent environmental exposure complicate both diagnosis and management. A critical clinical issue is determining whether recurrence represents true relapse, due to persistence of the original strain, or reinfection with a new strain, as this guides treatment and prevents overtreatment. Genotypic strategies capable of resolving strain-level differences can improve diagnostic accuracy, prevent misclassification, and ultimately support more informed treatment decisions. Therefore, integrating genotyping data into clinical workflows, standardizing single-nucleotide polymorphism thresholds, and establishing a global MAC strain database will not only support personalized treatment but also enhance the broader public health response to this disease.

Humans

Software for analysis and manipulation of genetic linkage data.

We present eight computer programs written in the C programming language that are designed to analyze genotypic data and to support existing software used to construct genetic linkage maps. Although each program has a unique purpose, they all share the common goals of affording a greater understanding of genetic linkage data and of automating tasks to make computers more effective tools for map building. The PIC/HET and FAMINFO programs automate calculation of relevant quantities such as heterozygosity, PIC, allele frequencies, and informativeness of markers and pedigrees. PREINPUT simplifies data submissions to the Centre d'Etude du Polymorphisme Humain (CEPH) data base by creating a file with genotype assignments that CEPH's INPUT program would otherwise require to be input manually. INHERIT is a program written specifically for mapping the X chromosome: by assigning a dummy allele to males, in the nonpseudoautosomal region, it eliminates falsely perceived noninheritances in the data set. The remaining four programs complement the previously published genetic linkage mapping software CRI-MAP and LINKAGE. TWOTABLE produces a more readable format for the output of CRI-MAP two-point calculations; UNMERGE is the converse to CRI-MAP's merge option; and GENLINK and LINKGEN automatically convert between the genotypic data file formats required by these packages. All eight applications read input from the same types of data files that are used by CRI-MAP and LINKAGE. Their use has simplified the management of data, has increased knowledge of the content of information in pedigrees, and has reduced the amount of time needed to construct genetic linkage maps of chromosomes.

Alleles

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Histologic, immunophenotypic and genotypic analyses of bone marrow trephines from patients with non-Hodgkin's lymphoma.

Marrow involvement in 20 patients with non-Hodgkin's lymphoma (NHL) were studied by histology, immunophenotypic and genotypic methods. Eighteen of these trephines were histologically involved with recognizable lymphomatous infiltrates and five of these were the primary disease site. In the remaining two cases (with histologically involved lymph nodes) the trephines were uninvolved with tumour. Three B-cell cases expressing surface immunoglobulin (sIg) and/or CD37 and one case not analysed phenotypically showed Ig gene rearrangements. The two remaining cases with B NHL showed no gene rearrangements, however, in one of these the trephine was histologically uninvolved with tumour. Twelve out of 14 T-cell cases were characterized by variable or absent expression of one or more T-cell antigens from the tumour population, one case was negative for all T-cell antigens and the remaining case was not histologically involved with tumour. All three lymphoblastic lymphomas and only 4/11 peripheral T-cell lymphomas (PTCL) cases revealed T-cell receptor (TcR) gene rearrangements. One of the latter cases also exhibited Ig JH gene rearrangements. This study demonstrates the usefulness of bone marrow trephines (BMT) in histologic, phenotypic and genotypic analyses. However, although genotypic data confirm clonality in B NHL and the lymphoblastic lymphomas there was genotypic heterogeneity within the PTCL group.

Antigens, CD

Genotypic identification of rickettsiae and estimation of intraspecies sequence divergence for portions of two rickettsial genes.

DNA sequences from specific genes, amplified by the polymerase chain reaction technique, were used as substrata for nonisotopic restriction endonuclease fragment length polymorphism differentiation of rickettsial species and genotypes. The products amplified using a single pair of oligonucleotide primers (derived from a rickettsial citrate synthase gene sequence) and cleaved with restriction endonucleases were used to differentiate almost all recognized species of rickettsiae. A second set of primers was used for differentiation of all recognized species of closely related spotted fever group rickettsiae. The procedure circumvents many technical obstacles previously associated with identification of rickettsial species. Multiple amplified DNA digest patterns were used to estimate the intraspecies nucleotide sequence divergence for the genes coding for rickettsial citrate synthase and a large antigen-coding gene of the spotted fever group rickettsiae. The estimated relationships deduced from these genotypic data correlate reasonably well with established rickettsial taxonomic schemes.

Antigens, Bacterial

Estimating the heritability of longitudinal rate-of-change: genetic insights into PSA velocity in prostate cancer-free individuals.

Serum prostate-specific antigen (PSA) is widely used for prostate cancer screening. While the genetics of PSA levels have been studied to enhance screening accuracy, the genetic basis of PSA velocity, the rate of PSA change over time, remains unclear. The Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial, a large, randomized study with longitudinal PSA data (15,260 cancer-free males, averaging 5.34 samples per subject) and genome-wide genotype data, provides a unique opportunity to estimate PSA velocity heritability. We developed a mixed model to jointly estimate the heritability of PSA levels at age 54 and PSA velocity. To accommodate the large dataset, we implemented 2 efficient computational approaches: a partitioning and meta-analysis strategy using average information restricted maximum likelihood (AI-REML) and a fast restricted Haseman-Elston (REHE) regression method. Simulations showed that both methods yield unbiased estimates of both heritability metrics, with AI-REML providing smaller variability in the estimation of velocity heritability than REHE. Applying AI-REML to PLCO data, we estimated heritability at 0.32 (s.e. = 0.07) for baseline PSA and 0.45 (s.e. = 0.18) for PSA velocity. These findings reveal a substantial genetic contribution to PSA velocity, supporting future genome-wide studies to identify variants affecting PSA dynamics and improve PSA-based screening.

Humans

Recall-by-genotype of neurodevelopmental disorder copy number variants in a multi-ancestry, healthcare-system biobank.

Clinical biobanks linking electronic health records (EHRs) with genotype data enable the study of genomic risk factors in real-world populations. However, recall-by-genotype (RbG) of psychiatric risk variants in diverse healthcare-system biobanks remains scarce. Leveraging BioMe, a multi-ancestry biobank within the Mount Sinai Health System, we recalled carriers of rare copy number variants (CNVs) that confer increased risk for neurodevelopmental disorders (NDDs) to establish empirical benchmarks for RbG implementation. We recontacted 892 participants: 335 NDD CNV carriers, 217 individuals with schizophrenia without NDD CNVs, and 340 neurotypical controls without NDD CNVs. Participants completed clinical and cognitive assessments. Overall, 18% of recontacted participants responded to recruitment, and 8% completed the study: 30 NDD CNV carriers, 20 individuals with schizophrenia, and 23 controls. The mean age was 48.8 years, 66% were female, and self-reported ancestry was 37% African, 34% Hispanic, and 26% European. Seventy percent of NDD CNV carriers had at least one neuropsychiatric or developmental condition, including mood or anxiety disorders (40%). Among 22 NDD CNV carriers at loci implicated in impaired cognition, performance was lower than controls on Digit Span Backward (&#x3b2;&#x2009;=&#x2009;-1.76, FDR&#x2009;=&#x2009;0.04) and Digit Span Sequencing (&#x3b2;&#x2009;=&#x2009;-2.01, FDR&#x2009;=&#x2009;0.04). NDD CNV carriers also outperformed the schizophrenia group on verbal learning (&#x3b2;&#x2009;=&#x2009;4.5, FDR&#x2009;=&#x2009;0.05). Recall of individuals-including those with psychiatric illness-yielded phenotypes not captured in EHRs and provides empirical benchmarks relevant to RbG implementation and precision psychiatry in diverse healthcare systems.

Journal Article