Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

A review of the 'Statistical Analysis for Genetic Epidemiology' (S.A.G.E.) software package.

The 'Statistical Analysis for Genetic Epidemiology' (S.A.G.E.) software package is an integrated, comprehensive package of computer programs designed to perform many of the different analyses required in the study of genetic epidemiology. It offers a graphical user interface for most platforms and, unlike many programs available in the public domain, is flexible in both receiving many types of input files and in allowing the user to choose among output files. All of the programs accept the same data files and together provide the means to perform familial correlation, segregation, linkage and association analyses, as well as many of the ancillary analyses that help achieve these goals. Many, but not all, of the same or similar analyses can be performed (with more difficulty) using publicly available freeware. The primary limitations of S.A.G.E. at present are the lack of software for estimating haplotypes or for identifying probable double recombinants in linkage analysis. S.A.G.E. is continually being extended and upgraded, however, with automatic downloading of the latest version always available to users.

Databases, Genetic↗

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation↗

A phylogenetic approach to assessing the significance of missense mutations in disease genes.

The identification of deleterious mutations within candidate genes is a crucial step in the elucidation of the genetic bases of human disease. However, the significance of any base or amino acid change within a gene is unknown until detailed structural and functional analysis has been carried out. A potentially rapid way of identifying functionally important sites within a gene is to identify evolutionarily conserved regions. Mutations affecting such sites are assumed to be deleterious for the carrier. In this communication we generalize this approach and present a formal framework to assess whether a specific mutation is deleterious given sequence data from a set of homologues. We propose a score that takes into account the nature of the mutation, the conservation of the affected residue among the different species, and their phylogenetic relationships. Its performance is examined using published TP53 mutations and frequent polymorphic variants.

Computational Biology↗

Genetics of Kidneys in Diabetes (GoKinD) study: a genetics collection available for identifying genetic susceptibility factors for diabetic nephropathy in type 1 diabetes.

The Genetics of Kidneys in Diabetes (GoKinD) study is an initiative that aims to identify genes that are involved in diabetic nephropathy. A large number of individuals with type 1 diabetes were screened to identify two subsets, one with clear-cut kidney disease and another with normal renal status despite long-term diabetes. Those who met additional entry criteria and consented to participate were enrolled. When possible, both parents also were enrolled to form family trios. As of November 2005, GoKinD included 3075 participants who comprise 671 case singletons, 623 control singletons, 272 case trios, and 323 control trios. Interested investigators may request the DNA collection and corresponding clinical data for GoKinD participants using the instructions and application form that are available at http://www.gokind.org/access. Participating scientists will have access to three data sets, each with distinct advantages. The set of 1294 singletons has adequate power to detect a wide range of genetic effects, even those of modest size. The set of case trios, which has adequate power to detect effects of moderate size, is not susceptible to false-positive results because of population substructure. The set of control trios is critical for excluding certain false-positive results that can occur in case trios and may be particularly useful for testing gene-environment interactions. Integration of the evidence from these three components into a single, unified analysis presents a challenge. This overview of the GoKinD study examines in detail the power of each study component and discusses analytic challenges that investigators will face in using this resource.

Adult↗

Software packages for quantitative microarray-based gene expression analysis.

Microarray technology enables researchers to investigate the expression of several thousand genes simultaneously. The whole transcriptional response of these genes in normal cells or tissue, in disease condition, as an response to biological, genetical or chemical stimuli or during normal biological processes such as cell cycle or embryonic development can be investigated. This leads to a huge amount of data, from which the relevant information has to be extracted by statistical and computational methods. Several software packages for the analysis of gene expression data are available, both commercially and freely. They differ particularly with regard to the implemented analytical methods, the graphical display and the manageability. In this paper the commercial software packages arraySCOUT, GeneSpring and Spotfire DecisionSite for Functional Genomics are compared and their applicability for analysis of gene expression data is studied. Small artificial and application test datasets are used to compare the computational results of the software packages. As far as possible results are verified with standard statistical software package SAS.

Algorithms↗

Genome Properties: a system for the investigation of prokaryotic genetic content for microbiology, genome annotation and comparative genomics.

MOTIVATION: The presence or absence of metabolic pathways and structures provide a context that makes protein annotation far more reliable. Compiling such information across microbial genomes improves the functional classification of proteins and provides a valuable resource for comparative genomics. RESULTS: We have created a Genome Properties system to present key aspects of prokaryotic biology using standardized computational methods and controlled vocabularies. Properties reflect gene content, phenotype, phylogeny and computational analyses. The results of searches using hidden Markov models allow many properties to be deduced automatically, especially for families of proteins (equivalogs) conserved in function since their last common ancestor. Additional properties are derived from curation, published reports and other forms of evidence. Genome Properties system was applied to 156 complete prokaryotic genomes, and is easily mined to find differences between species, correlations between metabolic features and families of uncharacterized proteins, or relationships among properties. AVAILABILITY: Genome Properties can be found at http://www.tigr.org/Genome_Properties SUPPLEMENTARY INFORMATION: http://www.tigr.org/tigr-scripts/CMR2/genome_properties_references.spl.

Chromosome Mapping↗

AACR Special Conference: SNPs, haplotypes, and cancer - applications in molecular epidemiology, Key Biscayne, Florida, USA, 13-17 September 2003.

This American Association for Cancer Research Special Conference brought together scientists with diverse expertise to address issues related to use of appropriate epidemiological, statistical, and laboratory methods to study the genetic epidemiology of cancer. Discussions focused on experiences with association studies using single nucleotide polymorphisms and haplotypes, their limitations, and what is needed to improve on the current 'state of the art'. Various studies were presented in different contexts, ranging from candidate gene studies to whole genome scans, and conducted in prospective cohorts, case-control studies, and other study designs. Common problems such as determining the probability that observed associations are false negative or false positive, the potential effects of admixture, and determining which polymorphisms to examine in which genes and in which populations were examined. Problems specific to haplotype analysis were discussed, with emphasis on haplotype block structures and on how to use haplotypes in analysis. Questions were also posed as to determining the functional relevance of single nucleotide polymorphisms in molecular epidemiology. Finally, future directions, using specific examples, were addressed.

DNA, Neoplasm↗

PyPop: a software framework for population genomics: analyzing large-scale multi-locus genotype data.

Software to analyze multi-locus genotype data for entire populations is useful for estimating haplotype frequencies, deviation from Hardy-Weinberg equilibrium and patterns of linkage disequilibrium. These statistical results are important to both those interested in human genome variation and disease predisposition as well as evolutionary genetics. As part of the 13th International Histocompatibility and Immunogenetics Working Group (IHWG), we have developed a software framework (PyPop). The primary novelty of this package is that it allows integration of statistics across large numbers of data-sets by heavily utilizing the XML file format and the R statistical package to view graphical output, while retaining the ability to inter-operate with existing software. Largely developed to address human population data, it can, however, be used for population based data for any organism. We tested our software on the data from the 13th IHWG which involved data sets from at least 50 laboratories each of up to 1000 individuals with 9 MHC loci (both class I and class II) and found that it scales to large numbers of data sets well.

Computational Biology↗

PowerTrim: An automated decision support algorithm for preprocessing family-based genetic data.

Statistical genetics software packages for linkage analysis have their own unique constraints on the size and shape of the pedigrees they can process. As a result, researchers are often forced to exclude from analysis some individuals in a given family. Existing procedures for reducing pedigree size to fit computational constraints use arbitrary rules and are not interactive. However, judicious evaluation of which subject(s) to remove to minimize loss of information involves consideration of many factors, including informativeness owing to position in pedigree, availability of genotypic information, and quality of phenotypic information. Thus, automation of this task would be of significant benefit. We designed an interactive algorithm (PowerTrim) that provides the user access to detailed information with which to make informed decisions. In addition, PowerTrim checks for transcriptional and data-entry errors, which can be very time-consuming to localize manually.

Algorithms↗

GRAST: a new way of genome reduction analysis using comparative genomics.

MOTIVATION: Establishment of intra-cellular life involved a profound re-configuration of the genetic characteristics of bacteria, including genome reduction and rearrangements. Understanding the mechanisms underlying these phenomena will shed light on the genome rearrangements essential for the development of an intra-cellular lifestyle. Comparison of genomes with differences in their sizes poses statistical as well as computational problems. Little efforts have been made to develop flexible computational tools with which to analyse genome reduction and rearrangements. RESULTS: Investigation of genome reduction and rearrangements in endosymbionts using a novel computational tool (GRAST) identified gathering of genes with similar functions. Conserved clusters of functionally related genes (CGSCs) were detected. Heterogeneous gene and gene cluster non-functionalization/loss are identified between genome regions, functional gene categories and during evolution. Results show that gene non-functionalisation has accelerated during the last 50 MY of Buchnera's evolution while CGSCs have been static.

Algorithms↗

Large-scale mutational analysis for the annotation of the mouse genome.

After sequencing the human and mouse genomes, the annotation of these sequences with biological functions is an important challenge in genomic research. A major tool to analyse gene function on the organismal level is the analysis of mutant phenotypes. Because of its genetic and physiological similarity to man, the mouse has become the model organism of choice for the study of genetic diseases. In addition, there is at the moment no other vertebrate for which versatile techniques to manipulate the genome are as well developed. Several mouse mutagenesis projects have provided the proof-of-principle that a systematic and comprehensive mutagenesis of every gene in the mammalian genome will be feasible. An exhaustive functional annotation of the mammalian genome can only be achieved in a combination of phenotype- and gene-driven approaches in large- and small-scale academic and private projects. Major challenges will be to develop standardised phenotyping protocols for the clinical and pathological characterisation of mouse mutants, the improvement of mutation detection methods and the dissemination of resources and data. Beyond gene annotation, it will be necessary to understand how gene functions are integrated into the complex network of regulatory interactions in the cell.

Animals↗

The role of informatics in the coordinated management of biological resources collections.

The term 'biological resources' is applied to the living biological material collected, held and catalogued in culture collections: bacterial and fungal cultures; animal, human and plant cells; viruses; and isolated genetic material. A wealth of information on these materials has been accumulated in culture collections, and most of this information is accessible. Digitalisation of data has reached a high level; however, information is still dispersed. Individual and coordinated approaches have been initiated to improve accessibility of biological resource centres, their holdings and related information through the Internet. These approaches cover subjects such as standardisation of data handling and data accessibility, and standardisation and quality control of laboratory procedures. This article reviews some of the most important initiatives implemented so far, as well as the most recent achievements. It also discusses the possible improvements that could be achieved by adopting new communication standards and technologies, such as web services, in view of a deeper and more fruitful integration of biological resources information in the bioinformatics network environment.

Animals↗

SNPselector: a web tool for selecting SNPs for genetic association studies.

SUMMARY: Single nucleotide polymorphisms (SNPs) are commonly used for association studies to find genes responsible for complex genetic diseases. With the recent advance of SNP technology, researchers are able to assay thousands of SNPs in a single experiment. But the process of manually choosing thousands of genotyping SNPs for tens or hundreds of genes is time consuming. We have developed a web-based program, SNPselector, to automate the process. SNPselector takes a list of gene names or a list of genomic regions as input and searches the Ensembl genes or genomic regions for available SNPs. It prioritizes these SNPs on their tagging for linkage disequilibrium, SNP allele frequencies and source, function, regulatory potential and repeat status. SNPselector outputs result in compressed Excel spreadsheet files for review by the user. AVAILABILITY: SNPselector is freely available at http://primer.duhs.duke.edu/

Algorithms↗

A scalable method for integration and functional analysis of multiple microarray datasets.

MOTIVATION: The diverse microarray datasets that have become available over the past several years represent a rich opportunity and challenge for biological data mining. Many supervised and unsupervised methods have been developed for the analysis of individual microarray datasets. However, integrated analysis of multiple datasets can provide a broader insight into genetic regulation of specific biological pathways under a variety of conditions. RESULTS: To aid in the analysis of such large compendia of microarray experiments, we present Microarray Experiment Functional Integration Technology (MEFIT), a scalable Bayesian framework for predicting functional relationships from integrated microarray datasets. Furthermore, MEFIT predicts these functional relationships within the context of specific biological processes. All results are provided in the context of one or more specific biological functions, which can be provided by a biologist or drawn automatically from catalogs such as the Gene Ontology (GO). Using MEFIT, we integrated 40 Saccharomyces cerevisiae microarray datasets spanning 712 unique conditions. In tests based on 110 biological functions drawn from the GO biological process ontology, MEFIT provided a 5% or greater performance increase for 54 functions, with a 5% or more decrease in performance in only two functions.

Algorithms↗

Mutation rates in mammalian genomes.

Knowledge of the rate of point mutation is of fundamental importance, because mutations are a vital source of genetic novelty and a significant cause of human diseases. Currently, mutation rate is thought to vary many fold among genes within a genome and among lineages in mammals. We have conducted a computational analysis of 5,669 genes (17,208 sequences) from species representing major groups of placental mammals to characterize the extent of mutation rate differences among genes in a genome and among diverse mammalian lineages. We find that mutation rate is approximately constant per year and largely similar among genes. Similarity of mutation rates among lineages with vastly different generation lengths and physiological attributes points to a much greater contribution of replication-independent mutational processes to the overall mutation rate. Our results suggest that the average mammalian genome mutation rate is 2.2 x 10(-9) per base pair per year, which provides further opportunities for estimating species and population divergence times by using molecular clocks.

Animals↗

Extension of multifactor dimensionality reduction for identifying multilocus effects in the GAW14 simulated data.

The multifactor dimensionality reduction (MDR) is a model-free approach that can identify gene x gene or gene x environment effects in a case-control study. Here we explore several modifications of the MDR method. We extended MDR to provide model selection without crossvalidation, and use a chi-square statistic as an alternative to prediction error (PE). We also modified the permutation test to provide different levels of stringency. The extended MDR (EMDR) includes three permutation tests (fixed, non-fixed, and omnibus) to obtain p-values of multilocus models. The goal of this study was to compare the different approaches implemented in the EMDR method and evaluate the ability to identify genetic effects in the Genetic Analysis Workshop 14 simulated data. We used three replicates from the simulated family data, generating matched pairs from family triads. The results showed: 1) chi-square and PE statistics give nearly consistent results; 2) results of EMDR without cross-validation matched that of EMDR with 10-fold cross-validation; 3) the fixed permutation test reports false-positive results in data from loci unrelated to the disease, but the non-fixed and omnibus permutation tests perform well in preventing false positives, with the omnibus test being the most conservative. We conclude that the non-cross-validation test can provide accurate results with the advantage of high efficiency compared to 10-cross-validation, and the non-fixed permutation test provides a good compromise between power and false-positive rate.

Computer Simulation↗

[Gene testing of hereditary cancer on comprehensive gene medical examination support system].

It has been estimated that genetic factors or a combination of genetic and environmental factors play a role in the development of 10-15% of all cancers. A genetic cause of hereditary cancer has been identified in more than 40 diseases till now. For preventing this cancer, gene testing is essential because it has no definite clinical marker as in hereditary non-polyposis colorectal cancer: HNPCC. Much more experience must be accumulated in this testing at the clinical base in order to increase specificity and sensitivity while safeguarding ethical, legal and social issues (ELSI). Recently, the Personal Information Protection Law was enforced. Gene inspection involving hereditary cancer should be carried out under a comprehensive gene medical examination organization. It is important for the family doctor, medical specialist, and gene inspection person in charge to cooperate closely with one another, and this will be a subject of future study.

Adenomatous Polyposis Coli↗

Haseman-Elston weighted by marker informativity.

In the Haseman-Elston approach the squared phenotypic difference is regressed on the proportion of alleles shared identical by descent (IBD) to map a quantitative trait to a genetic marker. In applications the IBD distribution is estimated and usually cannot be determined uniquely owing to incomplete marker information. At Genetic Analysis Workshop (GAW) 13, Jacobs et al. [BMC Genet 2003, 4(Suppl 1):S82] proposed to improve the power of the Haseman-Elston algorithm by weighting for information available from marker genotypes. The authors did not show, however, the validity of the employed asymptotic distribution. In this paper, we use the simulated data provided for GAW 14 and show that weighting Haseman-Elston by marker information results in increased type I error rates. Specifically, we demonstrate that the number of significant findings throughout the chromosome is significantly increased with weighting schemes. Furthermore, we show that the classical Haseman-Elston method keeps its nominal significance level when applied to the same data. We therefore recommend to use Haseman-Elston with marker informativity weights only in conjunction with empirical p-values. Whether this approach in fact yields an increase in power needs to be investigated further.

Chromosomes, Human, Pair 4↗