Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Comparative DNA sequence analysis of wheat and rice genomes.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to crop species has revolutionized molecular genetics and crop improvement strategies. This study compared 4485 expressed sequence tags (ESTs) that were physically mapped in wheat chromosome bins, to the public rice genome sequence data from 2251 ordered BAC/PAC clones using BLAST. A rice genome view of homologous wheat genome locations based on comparative sequence analysis revealed numerous chromosomal rearrangements that will significantly complicate the use of rice as a model for cross-species transfer of information in nonconserved regions.

Chromosome Mapping↗

Identifying genetic variation affecting a complex trait in simulated data: a comparison of meta-analysis with pooled data analysis.

We explored the power and consistency to detect linkage and association with meta-analysis and pooled data analysis using Genetic Analysis Workshop 14 simulated data. The first 10 replicates from Aipotu population were used. Significant linkage and association was found at all 4 regions containing the major loci for Kofendrerd Personality Disorder (KPD) using both combined analyses although no significant linkage and association was found at all these regions in a single replicate. The linkage results from both analyses are consistent in terms of the significance level of linkage test and the estimate of locus location. After correction for multiple-testing, significant associations were detected for the same 8 single-nucleotide polymorphisms (SNP) in both analyses. There were another 2 SNPs for which significant associations with KPD were found only by pooled data analysis. Our study showed that, under homogeneous condition, the results from meta-analysis and pooled data analysis are similar in both linkage and association studies and the loss of power is limited using meta-analysis. Thus, meta-analysis can provide an overall evaluation of linkage and association when the original raw data is not available for combining.

Computer Simulation↗

Recent applications of DNA microarray technology to toxicology and ecotoxicology.

Gene expression is a unique way of characterizing how cells and organisms adapt to changes in the external environment. The measurements of gene expression levels upon exposure to a chemical can be used both to provide information about the mechanism of action of the toxicant and to form a sort of "genetic signature" for the identification of toxic products. The development of high-quality, commercially available gene arrays has allowed this technology to become a standard tool in molecular toxicology. Several national and international initiatives have provided the proof-of-principle tests for the application of gene expression for the study of the toxicity of new and existing chemical compounds. In the last few years the field has progressed from evaluating the potential of the technology to illustrating the practical use of gene expression profiling in toxicology. The application of gene expression profiling to ecotoxicology is at an earlier stage, mainly because of the the many variables involved in analyzing the status of natural populations. Nevertheless, significant studies have been carried out on the response to environmental stressors both in model and in nonmodel organisms. It can be easily predicted that the development of stressor-specific signatures in gene expression profiling in ecotoxicology will have a major impact on the ecotoxicology field in the near future. International collaborations could play an important role in accelerating the application of genomic approaches in ecotoxicology.

Animals↗

Selection of minimum subsets of single nucleotide polymorphisms to capture haplotype block diversity.

We present a simple numerical algorithm to select the minimal subset of SNPs required to capture the diversity of haplotype blocks or other genetic loci. This algorithm can be used to quickly select the minimum SNP subset with no loss of haplotype information. In addition, the method can be used in a more aggressive mode to further reduce the original SNP set, with minimal loss of information. We demonstrate the algorithm performance with data from over 11,000 SNPs with average spacing of 6 to 11 Kb, across all the genes of chromosomes 6, 21, and 22, genotyped on DNA samples of 45 unrelated African-Americans and 45 Caucasians from the Coriell Human Diversity Collection. With no loss of information, we reduced the number of SNPs required to capture the haplotype block diversity by 25% for the African-American and 36% for the Caucasian populations. With a maximum loss of 10% of haplotype distribution information, the SNP reduction was 38% and 49% respectively for the two populations. All computations were performed in less than 1 minute for the entire dataset used.

Algorithms↗

Comparison of gene expression and DNA copy number changes in a murine model of lung cancer.

Activation of oncogenic Kras in murine lung leads to the development of numerous small adenomas, only some of which progress over time to overt adenocarcinoma. Thus, although Kras is the initiating oncogene, it is likely that secondary genetic events are required for progression from adenoma to adenocarcinoma. Some of these secondary events may also be important in human lung adenocarcinoma. By comparing gene expression profiles with DNA copy number changes, we sought to identify genes that play key roles in tumor progression in this model. Gene expression profiling revealed significant heterogeneity among the tumor samples. In 27% of the tumors analyzed, whole- or sub-chromosome duplications or deletions in one or more chromosomes were seen. Recurrent duplications were seen on chromosomes 6, 8, 16, and 19, whereas chromosomes 4, 11, and 17 were frequently lost. Notably, focal amplifications or deletions were not seen. Despite the lack of focal amplification, we showed that chromosome duplication has a measurable effect on gene expression that is not uniform across the genome. We identified a group of genes whose gene expression was highly correlated with changes in DNA copy number. These highly correlated genes were enriched for gene ontology categories involved in the DNA damage response and telomere maintenance.

Animals↗

A new protocol for evaluating putative causes for multiple variables in a spatial setting, illustrated by its application to European cancer rates.

We introduce a statistical protocol for analyzing spatially varying data, including putative explanatory variables. The procedures comprise preliminary spatial autocorrelation analysis (from an earlier study), path analysis, clustering of the resulting set of path diagrams, ordination of these diagrams, and confirmatory tests against extrinsic information. To illustrate the application of these methods, we present incidence and mortality rates of 31 organ- and sex-specific cancers in Europe; these rates vary markedly with geography and type of cancer. Additionally, we investigated three factors (ethnohistory, genetics, and geography) putatively affecting these rates. The five variables were correlated separately for the 31 cancers over European reporting stations. We analyzed the correlations by path analysis, k-means clustering, and nonmetric multidimensional scaling; coefficients of the 31 path diagrams modeling the correlations vary substantially. To simplify interpretation, we grouped the diagrams into five clusters, for which we describe the differential effects of the three putative causes on incidence and mortality. When scaled, the path coefficients intergrade without marked gaps between clusters. Ethnic differences make for differences in cancer rates, even when the populations tested are ancient and complex mixtures. Path analysis usefully decomposes a structural model involving effects and putative causes, and estimates the magnitude of the model's components. Smooth intergradation of the path coefficients suggests the putative causes are the results of multiple forces. Despite this continuity of the path diagrams of the 31 cancers, clustering offers a useful segmentation of the continuum. Etiological and other extrinsic information on the cancers map significantly into the five clusters, demonstrating their epidemiological relevance.

Biometry↗

OligoSpawn: a software tool for the design of overgo probes from large unigene datasets.

BACKGROUND: Expressed sequence tag (EST) datasets represent perhaps the largest collection of genetic information. ESTs can be exploited in a variety of biological experiments and analysis. Here we are interested in the design of overlapping oligonucleotide (overgo) probes from large unigene (EST-contigs) datasets. RESULTS: OLIGOSPAWN is a suite of software tools that offers two complementary services, namely (1) the selection of "unique" oligos each of which appears in one unigene but does not occur (exactly or approximately) in any other and (2) the selection of "popular" oligos each of which occurs (exactly or approximately) in as many unigenes as possible. In this paper, we describe the functionalities of OLIGOSPAWN and the computational methods it employs, and we report on experimental results for the overgo probes designed with it. CONCLUSION: The algorithms we designed are highly efficient and capable of processing unigene datasets of sizes on the order of several tens of Mb in a few hours on a regular PC. The software has been used to design overgo probes employed to screen a barley BAC library (Hordeum vulgare). OLIGOSPAWN is freely available at http://oligospawn.ucr.edu/.

Base Sequence↗

NMR structure of the p63 SAM domain and dynamical properties of G534V and T537P pathological mutants, identified in the AEC syndrome.

The p63 protein is crucial for epidermal development, and its mutations cause the extrodactyly ectodermal dysplasia and cleft lip/palate syndrome. The three-dimensional solution structure of the p63 sterile alpha-motif (SAM) domain (residues 505-579), a region crucial to explaining the human genetic disease ankyloblepharonectodermal dysplasia-clefting syndrome (AEC), has been determined by nuclear magnetic resonance spectroscopy. The structure indicates that the domain is a monomer with the characteristic five-helix bundle topology observed in other SAM domains. It includes five tightly packed helices with an extended hydrophobic core to form a globular and compact structure. The dynamics of the backbone and the global correlation time of the molecule have also been investigated and compared with the dynamical properties obtained through molecular dynamics simulation. Attempts to purify the pathological G534V and T537P mutants, originally identified in AEC, were not successful because of the occurrence of unspecific proteolytic degradation of the mutated SAM domains. Analysis of the structural dynamic properties of the G534V and T537P mutants through molecular dynamics simulation and comparison with the wild type permits detection of differences in the degree of freedom of individual residues and discussion of the possible causes for the pathology.

Abnormalities, Multiple↗

Clear relationship between ETF/ETFDH genotype and phenotype in patients with multiple acyl-CoA dehydrogenation deficiency.

Mutations in electron transfer flavoprotein (ETF) and its dehydrogenase (ETFDH) are the molecular basis of multiple acyl-CoA dehydrogenation deficiency (MADD), an autosomal recessively inherited and clinically heterogeneous disease that has been divided into three clinical forms: a neonatal-onset form with congenital anomalies (type I), a neonatal-onset form without congenital anomalies (type II), and a late-onset form (type III). To examine whether these different clinical forms could be explained by different ETF/ETFDH mutations that result in different levels of residual ETF/ETFDH enzyme activity, we have investigated the molecular genetic basis for disease development in nine patients representing the phenotypic spectrum of MADD. We report the genomic structures of the ETFA, ETFB, and ETFDH genes and the identification and characterization of seven novel and three previously reported disease-causing mutations. Our molecular genetic investigations of these nine patients are consistent with three clinical forms of MADD showing a clear relationship between the nature of the mutations and the severity of disease. Interestingly, our data suggest that homozygosity for two null mutations causes fetal development of congenital anomalies resulting in a type I disease phenotype. Even minute amounts of residual ETF/ETFDH activity seem to be sufficient to prevent embryonic development of congenital anomalies giving rise to type II disease. Overexpression studies of an ETFB-D128N missense mutation identified in a patient with type III disease showed that the residual activity of the mutant enzyme could be rescued up to 59% of that of wild-type activity when ETFB-D128N-transformed E. coli cells were grown at low temperature. This indicates that the effect of the ETF/ETFDH genotype in patients with milder forms of MADD, in whom residual enzyme activity allows modulation of the enzymatic phenotype, may be influenced by environmental factors like cellular temperature.

Acyl-CoA Dehydrogenase↗

Grass evolution inferred from chromosomal rearrangements and geometrical and statistical features in RNA structure.

The grasses (Poaceae) represent a monophyletic lineage that arose about 70 million years ago. The lineage contains about 10,000 species that differ widely in morphology and physiology. Species show striking differences in genome size, a feature important in the context of conservation of gene content and order (synteny and colinearity) and in the extension of genomic information directly from one grass species to another using comparative approaches. Grass diversification has been a contentious issue, as the exact branching order of the various subfamilies has been difficult to establish with standard methods. This motivated an evolutionary study of deep phylogenetic relationships based on the structure of coding and non-coding RNA molecules and on chromosomal rearrangements. Phylogenetic relationships in the grass family were inferred directly from the structure of RNA using cladistic principles and considerations in statistical mechanics. Coded attributes describing topological and thermodynamic information embedded in RNA molecules were treated as linearly ordered multi-state characters and were polarized by fixing the direction of character transformation toward molecular order. Intrinsically rooted phylogenies derived from the structure of signal recognition particle (SRP) RNA, the mRNA encoded by the early nodulation gene enod40, the small subunit of ribosomal RNA (rRNA), and the internal transcribed spacer ITS1 of rRNA established an order for the diversification of major grass lineages, suggesting a sister relationship of the Pooideae and the PACCAD clade. This same conclusion was reached when large-scale chromosomal rearrangements derived from the comparative genetic mapping of cereal genomes were studied. Chromosomal complements aligned in the most parsimonious manner allowed identification and coding of characters depicting chromosomal translocations, insertions, and linkage block arrangements and the reconstruction of phylogenetic trees based on large-scale chromosomal structure. Congruent reconstruction of deep branching relationships using geometrical and statistical features of RNA structure and orthology and large scale chromosomal recombination events support assumptions of polarization in character argumentation, and fail to falsify the claim that extant grass chromosomes can be considered combinations of linkage blocks of an ancestor of the rice genome. Congruence also suggests that the universal tendency toward order in RNA and the search for the most parsimonious organization of be genome architecture appear to be mutually supported drivers of molecular evolution. The study clarifies the relationship of major clades in the grasses, shows that phylogenetic history can be reconstructed effectively from the combinatorial exchange of chromosomal linkage blocks, and reveals considerable phylogenetic signal embedded in the structure of signal polypeptide-coding mRNA molecules, describing an instance where mRNA structure is the subject of strong evolutionary constraint.

Base Pairing↗

Automated SNP detection from a large collection of white spruce expressed sequences: contributing factors and approaches for the categorization of SNPs.

BACKGROUND: High-throughput genotyping technologies represent a highly efficient way to accelerate genetic mapping and enable association studies. As a first step toward this goal, we aimed to develop a resource of candidate Single Nucleotide Polymorphisms (SNP) in white spruce (Picea glauca [Moench] Voss), a softwood tree of major economic importance. RESULTS: A white spruce SNP resource encompassing 12,264 SNPs was constructed from a set of 6,459 contigs derived from Expressed Sequence Tags (EST) and by using the bayesian-based statistical software PolyBayes. Several parameters influencing the SNP prediction were analysed including the a priori expected polymorphism, the probability score (PSNP), and the contig depth and length. SNP detection in 3' and 5' reads from the same clones revealed a level of inconsistency between overlapping sequences as low as 1%. A subset of 245 predicted SNPs were verified through the independent resequencing of genomic DNA of a genotype also used to prepare cDNA libraries. The validation rate reached a maximum of 85% for SNPs predicted with either PSNP > or = 0.95 or > or = 0.99. A total of 9,310 SNPs were detected by using PSNP > or = 0.95 as a criterion. The SNPs were distributed among 3,590 contigs encompassing an array of broad functional categories, with an overall frequency of 1 SNP per 700 nucleotide sites. Experimental and statistical approaches were used to evaluate the proportion of paralogous SNPs, with estimates in the range of 8 to 12%. The 3,789 coding SNPs identified through coding region annotation and ORF prediction, were distributed into 39% nonsynonymous and 61% synonymous substitutions. Overall, there were 0.9 SNP per 1,000 nonsynonymous sites and 5.2 SNPs per 1,000 synonymous sites, for a genome-wide nonsynonymous to synonymous substitution rate ratio (Ka/Ks) of 0.17. CONCLUSION: We integrated the SNP data in the ForestTreeDB database along with functional annotations to provide a tool facilitating the choice of candidate genes for mapping purposes or association studies.

Algorithms↗

Combinatorial microarray analysis revealing arabidopsis genes implicated in cytokinin responses through the His->Asp Phosphorelay circuitry.

In Arabidopsis thaliana, the immediate early response of plants to cytokinin is formulated as the multistep histidine kinase (AHK)-->histidine-containing phosphotransmitter (AHP)-->response regulator (ARR) phosphorelay signaling circuitry, which is initiated by the cytokinin receptor histidine protein kinases. In the hope of finding components (or genes) that function downstream of the cytokinin-mediated His-->Asp phosphorelay signaling circuitry, we carried out genome-wide microarray analyses. To this end, we used a combinatorial microarray strategy by employing not only wild-type plants, but also certain transgenic lines in which the cytokinin-mediated His-->Asp phosphorelay signaling circuitry has been genetically manipulated. These transgenic lines employed were ARR21-overexpressing and ARR22-overexpressing plants, each of which exhibits a characteristic phenotype with regard to the cytokinin-mediated His-->Asp phosphorelay. The results of extensive microarray analyses with these plants allowed us systematically to identify a certain number of genes that were up-regulated at the level of transcription in response to cytokinin directly or indirectly. Among them, some representatives were examined further in wild-type plants to support the idea that certain genes encoding transcription factors are rapidly and specifically induced at the level of transcription by cytokinin in a manner similar to that of the type-A ARR genes, which are the hallmarks of the His-->Asp phosphorelay signaling circuitry. Several interesting transcription factors were thus identified as being cytokinin responsive, including those belonging to the AP2/EREBP family, MYB family, GATA family or bHLH family. Including these, the presented list of cytokinin-up-regulated genes (214) will provide us with valuable bases for understanding the His-->Asp phosphorelay in A. thaliana.

Arabidopsis↗

Physiogenomic resources for rat models of heart, lung and blood disorders.

Cardiovascular disorders are influenced by genetic and environmental factors. The TIGR rodent expression web-based resource (TREX) contains over 2,200 microarray hybridizations, involving over 800 animals from 18 different rat strains. These strains comprise genetically diverse parental animals and a panel of chromosomal substitution strains derived by introgressing individual chromosomes from normotensive Brown Norway (BN/NHsdMcwi) rats into the background of Dahl salt sensitive (SS/JrHsdMcwi) rats. The profiles document gene-expression changes in both genders, four tissues (heart, lung, liver, kidney) and two environmental conditions (normoxia, hypoxia). This translates into almost 400 high-quality direct comparisons (not including replicates) and over 100,000 pairwise comparisons. As each individual chromosomal substitution strain represents on average less than a 5% change from the parental genome, consomic strains provide a useful mechanism to dissect complex traits and identify causative genes. We performed a variety of data-mining manipulations on the profiles and used complementary physiological data from the PhysGen resource to demonstrate how TREX can be used by the cardiovascular community for hypothesis generation.

Animals↗

Defining synphenotype groups in Xenopus tropicalis by use of antisense morpholino oligonucleotides.

To identify novel genes involved in early development, and as proof-of-principle of a large-scale reverse genetics approach in a vertebrate embryo, we have carried out an antisense morpholino oligonucleotide (MO) screen in Xenopus tropicalis, in the course of which we have targeted 202 genes expressed during gastrula stages. MOs were designed to complement sequence between -80 and +25 bases of the initiating AUG codons of the target mRNAs, and the specificities of many were tested by (i) designing different non-overlapping MOs directed against the same mRNA, (ii) injecting MOs differing in five bases, and (iii) performing "rescue" experiments. About 65% of the MOs caused X. tropicalis embryos to develop abnormally (59% of those targeted against novel genes), and we have divided the genes into "synphenotype groups," members of which cause similar loss-of-function phenotypes and that may function in the same developmental pathways. Analysis of the expression patterns of the 202 genes indicates that members of a synphenotype group are not necessarily members of the same synexpression group. This screen provides new insights into early vertebrate development and paves the way for a more comprehensive MO-based analysis of gene function in X. tropicalis.

Animals↗

Data mining of the GAW14 simulated data using rough set theory and tree-based methods.

Rough set theory and decision trees are data mining methods used for dealing with vagueness and uncertainty. They have been utilized to unearth hidden patterns in complicated datasets collected for industrial processes. The Genetic Analysis Workshop 14 simulated data were generated using a system that implemented multiple correlations among four consequential layers of genetic data (disease-related loci, endophenotypes, phenotypes, and one disease trait). When information of one layer was blocked and uncertainty was created in the correlations among these layers, the correlation between the first and last layers (susceptibility genes and the disease trait in this case), was not easily directly detected. In this study, we proposed a two-stage process that applied rough set theory and decision trees to identify genes susceptible to the disease trait. During the first stage, based on phenotypes of subjects and their parents, decision trees were built to predict trait values. Phenotypes retained in the decision trees were then advanced to the second stage, where rough set theory was applied to discover the minimal subsets of genes associated with the disease trait. For comparison, decision trees were also constructed to map susceptible genes during the second stage. Our results showed that the decision trees of the first stage had accuracy rates of about 99% in predicting the disease trait. The decision trees and rough set theory failed to identify the true disease-related loci.

Computer Simulation↗

Ectrodactyly with aplasia of long bones (OMIM; 119100) in a large inbred Arab family with an apparent autosomal dominant inheritance and reduced penetrance: clinical and genetic analysis.

Ectrodactyly with aplasia of long bones syndrome is one of the most recognizable defects involving the extremities. We have studied a very large eight-generation consanguineous Arab family from the United Arab Emirates (UAE) with multiple severe limb anomalies resembling this condition (OMIM; 119100), for which the affected gene is unknown. The pedigree consists of 145 individuals including 23 affected (14 males/9 females) with limb anomalies. Of these, 18 had tibial aplasia (TA) usually on the right side. The expression of the phenotype was variable and ranged from bilateral to unilateral TA with ectrodactyly and other defects of the extremities. The mode of inheritance appears to be autosomal dominant with reduced penetrance. There were 10 consanguineous marriages observed in this pedigree. This could suggest possible pseudodominance due to high frequency of the mutant allele. Candidate loci for the described syndrome include GLI3 (OMIM: 165240) on 7p13, sonic hedgehog; (OMIM: 600725) on 7q36, Langer-Giedion syndrome (OMIM: 150230) on 8q24.1 and split-hand/foot malformation 3 (OMIM: 600095) on 10q24. In addition, bilateral tibial hemimelia and unilateral absence of the ulna was previously observed to co-segregate with deletion of 8q24.1. Two-point linkage and haplotype analyses did not show the involvement of the above regions in this family.

Abnormalities, Multiple↗

Navigating the HapMap.

With the availability of the HapMap--a resource which describes common patterns of linkage disequilibrium (LD) in four different human population samples, we now have a powerful tool to help dissect the role of genetic variation in the biology of the genome. HapMap is entirely complimentary to the human genome map and so it is particularly fitting that it should be viewed in a full genomic context. However, characterization of high resolution LD across the genome can be a challenging task, owing in part to the sheer volume of data and the inherent dimensionality that its analysis entails. However, a number of tools are now available to make this task easier for researchers. This review will examine tools for viewing and analysing haplotype and LD data, enabling a number of tasks; including identification of optimal sets of haplotype tagging single nucleotide polymorphisms (SNPs); drawing links between associated SNPs and putative causal alleles; or simply viewing LD and haplotypes across a gene or region of interest. The data generated by the HapMap also has other important applications, informing, for example, on the demographic history and evidence of selection in human populations and on previously undetected regulatory relationships and gene networks. All of these properties make the HapMap no less an important resource than the human genome sequence itself and so this makes it essential viewing for all in the field of human biology.

Alleles↗