Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Comparative analysis of gene-expression patterns in human and African great ape cultured fibroblasts.

Although much is known about genetic variation in human and African great ape (chimpanzee, bonobo, and gorilla) genomes, substantially less is known about variation in gene-expression profiles within and among these species. This information is necessary for defining transcriptional regulatory networks that contribute to complex phenotypes unique to humans or the African great apes. We took a systematic approach to this problem by investigating gene-expression profiles in well-defined cell populations from humans, bonobos, and gorillas. By comparing these profiles from 18 human and 21 African great ape primary fibroblast cell lines, we found that gene-expression patterns could predict the species, but not the age, of the fibroblast donor. Several differentially expressed genes among human and African great ape fibroblasts involved the extracellular matrix, metabolic pathways, signal transduction, stress responses, as well as inherited overgrowth and neurological disorders. These gene-expression patterns could represent molecular adaptations that influenced the development of species-specific traits in humans and the African great apes.

Africa↗

Cloning and characterization of three PHEX homologues in Drosophila.

Inactivating mutations and/or deletions of PHEX ( Phosphate-regulating gene with Homologies to Endopeptidase on the X chromosome) are responsible for X-linked hypophosphatemic rickets in humans. In the present study, three Drosophila PHEX homologues (dPHEX-1, -2, -3) were isolated by the screening of a Drosophila cDNA library and expressed sequence tag (EST) database. The structural region involving motif II: (456)WMXXXTKXXAXXK(468) (numbered according to human PHEX), motif VI: (602)WW(603), and motif VIII: (746)CXLW(749) was conserved in the dPHEX family. Zinc-coordinating motifs (HEFTH and GENIADNGG) were also conserved in the dPHEX family. All three dPHEX genes were expressed during all stages of Drosophila development. The expression of dPHEX-1 was suppressed by dietary phosphate deprivation, but the expression of dPHEX-2 and that of dPHEX-3 were not affected. In-situ hybridization showed a ubiquitous distribution of dPHEX-1 and dPHEX-2, while dPHEX-3 was highly expressed in the larval brain. In an analysis of subcellular localization, dPHEX-1 was localized to intracellular organelles and dPHEX-3 was localized predominately in the plasma membrane of Drosophila embryonic S2 cells. Homozygosity of a dPHEX-1 mutation, a transposon insertion in the dPHEX-1 promoter region, was completely lethal at an early stage of embryonic development. The present study indicates that three homologues are likely involved in the phosphate homeostasis of Drosophila.

Alternative Splicing↗

IMGT unique numbering for MHC groove G-DOMAIN and MHC superfamily (MhcSF) G-LIKE-DOMAIN.

IMGT, the international ImMunoGeneTics information system (http://imgt.cines.fr) provides a common access to expertly annotated data on the genome, proteome, genetics and structure of immunoglobulins (IG), T cell receptors (TR), major histocompatibility complex (MHC), and related proteins of the immune system (RPI) of human and other vertebrates. The NUMEROTATION concept of IMGT-ONTOLOGY has allowed to define a unique numbering for the variable domains (V-DOMAINs) and constant domains (C-DOMAINs) of the IG and TR, which has been extended to the V-LIKE-DOMAINs and C-LIKE-DOMAINs of the immunoglobulin superfamily (IgSF) proteins other than the IG and TR (Dev Comp Immunol 27:55--77, 2003; 29:185--203, 2005). In this paper, we describe the IMGT unique numbering for the groove domains (G-DOMAINs) of the MHC and for the G-LIKE-DOMAINs of the MHC superfamily (MhcSF) proteins other than MHC. This IMGT unique numbering leads, for the first time, to the standardized description of the mutations, allelic polymorphisms, two-dimensional (2D) representations and three-dimensional (3D) structures of the G-DOMAINs and G-LIKE-DOMAINs in any species, and therefore, is highly valuable for their comparative, structural, functional and evolutionary studies.

Amino Acid Sequence↗

Genomic research and data-mining technology: implications for personal privacy and informed consent.

This essay examines issues involving personal privacy and informed consent that arise at the intersection of information and communication technology (ICT) and population genomics research. I begin by briefly examining the ethical, legal, and social implications (ELSI) program requirements that were established to guide researchers working on the Human Genome Project (HGP). Next I consider a case illustration involving deCODE Genetics, a privately owned genetic company in Iceland, which raises some ethical concerns that are not clearly addressed in the current ELSI guidelines. The deCODE case also illustrates some ways in which an ICT technique known as data mining has both aided and posed special challenges for researchers working in the field of population genomics. On the one hand, data-mining tools have greatly assisted researchers in mapping the human genome and in identifying certain "disease genes" common in specific populations (which, in turn, has accelerated the process of finding cures for diseases tha affect those populations). On the other hand, this technology has significantly threatened the privacy of research subjects participating in population genomics studies, who may, unwittingly, contribute to the construction of new groups (based on arbitrary and non-obvious patterns and statistical correlations) that put those subjects at risk for discrimination and stigmatization. In the final section of this paper I examine some ways in which the use of data mining in the context of population genomics research poses a critical challenge for the principle of informed consent, which traditionally has played a central role in protecting the privacy interests of research subjects participating in epidemiological studies.

Computational Biology↗

CYP2A6 polymorphisms in Malays, Chinese and Indians.

The genetically polymorphic cytochrome P450 (CYP) 2A6 is the major nicotine-oxidase in humans that may contribute to nicotine dependence and cancer susceptibility. The authors investigated the types and frequencies of CYP2A6 alleles in the three major ethnic groups in Malaysia and CYP2A6*1A, CYP2A6*1B, CYP2A6*1x2, CYP2A6*2, CYP2A6*3, CYP2A6*4, CYP2A6*5, CYP2A6*7, CYP2A6*8 and CYP2A6*10 were determined by allele-specific polymerase chain reaction (PCR) in 270 Malays, 172 Chinese and 174 Indians. Except for CYP2A6*2 and *3 that were not detected in the Malays and Chinese, all the other alleles were detected. Frequencies for the CYP2A6*4 allele were 7, 5 and 2%, respectively, in Malays, Chinese and Indians. A statistically significant high frequency of the duplicated CYP2A6*1x2 allele occurred among Chinese. Among Malays and Chinese, the most common allele was CYP2A6*1B, but it was CYP2A6*1A among Indians. These ethnic difference in frequencies suggested that further studies are required to investigate the implications on diseases such as cancer and smoking behaviour among these major ethnic groups in Malaysia.

Adult↗

A chromosomal rearrangement hotspot can be identified from population genetic variation and is coincident with a hotspot for allelic recombination.

Insights into the origins of structural variation and the mutational mechanisms underlying genomic disorders would be greatly improved by a genomewide map of hotspots of nonallelic homologous recombination (NAHR). Moreover, our understanding of sequence variation within the duplicated sequences that are substrates for NAHR lags far behind that of sequence variation within the single-copy portion of the genome. Perhaps the best-characterized NAHR hotspot lies within the 24-kb-long Charcot-Marie-Tooth disease type 1A (CMT1A)-repeats (REPs) that sponsor deletions and duplications that cause peripheral neuropathies. We investigated structural and sequence diversity within the CMT1A-REPs, both within and between species. We discovered a high frequency of retroelement insertions, accelerated sequence evolution after duplication, extensive paralogous gene conversion, and a greater than twofold enrichment of SNPs in humans relative to the genome average. We identified an allelic recombination hotspot underlying the known NAHR hotspot, which suggests that the two processes are intimately related. Finally, we used our data to develop a novel method for inferring the location of an NAHR hotspot from sequence variation within segmental duplications and applied it to identify a putative NAHR hotspot within the LCR22 repeats that sponsor velocardiofacial syndrome deletions. We propose that a large-scale project to map sequence variation within segmental duplications would reveal a wealth of novel chromosomal-rearrangement hotspots.

Alleles↗

Mining genetic epidemiology data with Bayesian networks I: Bayesian networks and example application (plasma apoE levels).

MOTIVATION: The wealth of single nucleotide polymorphism (SNP) data within candidate genes and anticipated across the genome poses enormous analytical problems for studies of genotype-to-phenotype relationships, and modern data mining methods may be particularly well suited to meet the swelling challenges. In this paper, we introduce the method of Belief (Bayesian) networks to the domain of genotype-to-phenotype analyses and provide an example application. RESULTS: A Belief network is a graphical model of a probabilistic nature that represents a joint multivariate probability distribution and reflects conditional independences between variables. Given the data, optimal network topology can be estimated with the assistance of heuristic search algorithms and scoring criteria. Statistical significance of edge strengths can be evaluated using Bayesian methods and bootstrapping. As an example application, the method of Belief networks was applied to 20 SNPs in the apolipoprotein (apo) E gene and plasma apoE levels in a sample of 702 individuals from Jackson, MS. Plasma apoE level was the primary target variable. These analyses indicate that the edge between SNP 4075, coding for the well-known epsilon2 allele, and plasma apoE level was strong. Belief networks can effectively describe complex uncertain processes and can both learn from data and incorporate prior knowledge. AVAILABILITY: Various alternative and supplemental networks (not given in the text) as well as source code extensions, are available from the authors. SUPPLEMENTARY INFORMATION: http://bioinformatics.oxfordjournals.org.

Apolipoproteins E↗

From phenotype to genotype: issues in navigating the available information resources.

OBJECTIVES: As part of an investigation of connecting health professionals and the lay public to both disease and genomic information, we assessed the availability and nature of the data from the Human Genome Project relating to human genetic diseases. METHODS: We focused on a set of single gene diseases selected from main topics in MEDLINEplus, the NLM's principal resource focused on consumers. We used publicly available websites to investigate specific questions about the genes and gene products associated with the diseases. We also investigated questions of knowledge and data representation for the information resources and navigational issues. RESULTS: Many online resources are available but they are complex and technical. The major challenges encountered when navigating from phenotype to genotype were (1) complexity of the data, (2) dynamic nature of the data, (3) diversity of foci and number of information resources, and (4) lack of use of standard data and knowledge representation methods. CONCLUSIONS: Three major informatics issues arise from the navigational challenges. First, the official gene names are insufficient for navigation of these web resources. Second, navigational inconsistencies arise from difficulties in determining the number and function of alternate forms of the gene or gene product and maintaining currency with this information. Third, synonymy and polysemy cause much confusion. These are severe obstacles to computational navigation from phenotype to genotype, especially for individuals who are novices in the underlying science. Tools and standards to facilitate this navigation are sorely needed.

Databases, Genetic↗

Heterogeneity-based genome search meta-analysis for preeclampsia.

Preeclampsia is a pregnancy-related disorder that causes maternal and fetal morbidity and mortality. Its exact inheritance pattern is still unknown, and genome searches for identifying susceptibility loci for preeclampsia have thus far produced inconclusive or inconsistent results. We performed a heterogeneity-based genome search meta-analysis (HEGESMA) that synthesized the available genome scan data on preeclampsia. HEGESMA identifies genetic regions (bins) that rank highly on average in terms of linkage statistics across genome scans (searches). The significance of each bin's average rank and heterogeneity across scans was calculated using Monte Carlo tests. The meta-analysis involved four genome-scans on general preeclampsia and five scans on severe preeclampsia. In general preeclampsia, 13 bins had significantly high average rank (Prank< 0.05) by either unweighted or weighted analyses, while four of them (2p11.2-2q21.1, 9q21.32-9q31.2, 2p15-2p11.2, 2q32.1-2q35) were formally significant by both analyses. Heterogeneity of bin 2.8 (2q32.1-2q35) was significantly low in both unweighted and weighted analysis (PQ< 0.01). In severe preeclampsia, 10 bins had significantly high average rank by either unweighted or weighted analyses and five of them (3q11.1-3q21.2, 2q37.1-2q37.3, 18p11.32-18p11.22, 2p15-2p11.2, 7q34-7q36.3) were significant by both analyses. Bin 2q37.1-2q37.3 showed marginal low heterogeneity in unweighted and weighted analysis (PQ= 0.06). Results should be interpreted with caution as the p values were modest. Further investigation of these regions by genotyping with additional markers and families may help to direct the identification of candidate genes for preeclampsia.

Chromosome Mapping↗

Advanced integrated mouse YAC map including BAC framework.

Functional characterization of the mouse genome requires the availability of a comprehensive physical map to obtain molecular access to chromosomal regions of interest. Positional cloning remains a crucial way of linking phenotype with particular genes. A key step and frequent stumbling block in positional cloning is making a contig of a genetically defined candidate region. The most efficient first step is isolating YAC (Yeast Artificial Chromosome) clones. A robust, detailed YAC contig map is thus an important tool. Employing Interspersed Repetitive Sequence (IRS)-PCR genomics, we have generated an advanced second-generation YAC contig map of the mouse genome that doubles both the depth of clones and the density of markers available. In addition to the primarily YAC-based map, we located 1942 BAC (Bacterial Artificial Chromosome) clones. This allows us to present for the first time a dense framework of BACs spanning the genome of the mouse, which, for instance, can serve as a nucleus for genomic sequencing. Four large-insert mouse YAC libraries from three different strains are included in our data, and our analysis incorporates the data of Hunter et al. and Nusbaum et al. There is a total of 20,205 markers on the final map, 12,033 from our own data, and a total of 56,093 YACs, of which 44,401 are positive for more than one marker.

Algorithms↗

A genomewide survey of developmentally relevant genes in Ciona intestinalis. IX. Genes for muscle structural proteins.

Ascidians are simple chordates that are related to, and may resemble, vertebrate ancestors. Comparison of ascidian and vertebrate genomes is expected to provide insight into the molecular genetic basis of chordate/vertebrate evolution. We annotated muscle structural (contractile protein) genes in the completely determined genome sequence of the ascidian Ciona intestinalis, and examined gene expression patterns through extensive EST analysis. Ascidian muscle protein isoform families are generally of similar, or lesser, complexity in comparison with the corresponding vertebrate isoform families, and are based on gene duplication histories and alternative splicing mechanisms that are largely or entirely distinct from those responsible for generating the vertebrate isoforms. Although each of the three ascidian muscle types - larval tail muscle, adult body-wall muscle and heart - expresses a distinct profile of contractile protein isoforms, none of these isoforms are strictly orthologous to the smooth-muscle-specific, fast or slow skeletal muscle-specific, or heart-specific isoforms of vertebrates. Many isoform families showed larval-versus-adult differential expression and in several cases numerous very similar genes were expressed specifically in larval muscle. This may reflect different functional requirements of the locomotor larval muscle as opposed to the non-locomotor muscles of the sessile adult, and/or the biosynthetic demands of extremely rapid larval development.

Amino Acid Sequence↗

A novel method for automatic genotyping of microsatellite markers based on parametric pattern recognition.

Genetic mapping of loci affecting complex phenotypes in human and other organisms is presently being conducted on a very large scale, using either microsatellite or single nucleotide polymorphism (SNP) markers and by partly automated methods. A critical step in this process is the conversion of the instrument output into genotypes, both a time-consuming and error prone procedure. Errors made during this calling of genotypes will dramatically reduce the ability to map the location of loci underlying a phenotype. Accurate methods for automatic genotype calling are therefore important. Here, we describe novel algorithms for automatic calling of microsatellite genotypes using parametric pattern recognition. The analysis of microsatellite data is complicated both by the occurrence of stutter bands, which arise from Taq polymerase misreading the number of repeats, and additional bands derived form the non-template dependent addition of a nucleotide to the 3' end of the PCR products. These problems, together with the fact that the lengths of two alleles in a heterozygous individual may differ by only two nucleotides, complicate the development of an automated process. The novel algorithms markedly reduce the need for manual editing and the frequency of miscalls, and compares very favourably with commercially available software for automatic microsatellite genotyping.

Algorithms↗

Application of DETECTER, an evolutionary genomic tool to analyze genetic variation, to the cystic fibrosis gene family.

BACKGROUND: The medical community requires computational tools that distinguish missense genetic differences having phenotypic impact within the vast number of sense mutations that do not. Tools that do this will become increasingly important for those seeking to use human genome sequence data to predict disease, make prognoses, and customize therapy to individual patients. RESULTS: An approach, termed DETECTER, is proposed to identify sites in a protein sequence where amino acid replacements are likely to have a significant effect on phenotype, including causing genetic disease. This approach uses a model-dependent tool to estimate the normalized replacement rate at individual sites in a protein sequence, based on a history of those sites extracted from an evolutionary analysis of the corresponding protein family. This tool identifies sites that have higher-than-average, average, or lower-than-average rates of change in the lineage leading to the sequence in the population of interest. The rates are then combined with sequence data to determine the likelihoods that particular amino acids were present at individual sites in the evolutionary history of the gene family. These likelihoods are used to predict whether any specific amino acid replacements, if introduced at the site in a modern human population, would have a significant impact on fitness. The DETECTER tool is used to analyze the cystic fibrosis transmembrane conductance regulator (CFTR) gene family. CONCLUSION: In this system, DETECTER retrodicts amino acid replacements associated with the cystic fibrosis disease with greater accuracy than alternative approaches. While this result validates this approach for this particular family of proteins only, the approach may be applicable to the analysis of polymorphisms generally, including SNPs in a human population.

Amino Acid Substitution↗

Evaluation of an algorithm of tagging SNPs selection by linkage disequilibrium.

BACKGROUND: Single nucleotide polymorphisms (SNPs) are the most abundant kind of genetic polymorphism in the human genome. They are important in both genetic research and genetic testing in a clinical setting, such as in the area of pharmacogenetics. In order to improve efficiency, tagging SNPs (tagSNPs) are selected in genes of interest to represent other co-related SNPs in linkage disequilibrium (LD) with the tagSNPs. Various algorithms have been proposed to identify a subset of single nucleotide polymorphisms as tagSNPs. Most algorithms of tagSNPs selection are haplotype-based, in which the spatial relationship between SNPs is considered. Currently, a more efficient cluster-based algorithm is proposed which clusters SNPs solely by a LD parameter, such as r(2). Here, we evaluated the sample distribution of r(2) and its effect on the cluster-based tagSNPs selection. DESIGN AND METHODS: The genotype data of 198 individual within a 500-kb region on 5q31 was used to evaluate the sample distribution of r(2) and its effect on the cluster-based tagSNPs selection. RESULTS: It was found that the degree of variation of LD depends on the LD structure of genes. CONCLUSION: As a cluster-based tagSNPs selection algorithm does not take into account the spatial position of SNPs, a more stringent r(2) threshold is required to achieve more reliable tagSNPs selection.

Algorithms↗

Analysis of the genome-wide variations among multiple strains of the plant pathogenic bacterium Xylella fastidiosa.

BACKGROUND: The Gram-negative, xylem-limited phytopathogenic bacterium Xylella fastidiosa is responsible for causing economically important diseases in grapevine, citrus and many other plant species. Despite its economic impact, relatively little is known about the genomic variations among strains isolated from different hosts and their influence on the population genetics of this pathogen. With the availability of genome sequence information for four strains, it is now possible to perform genome-wide analyses to identify and categorize such DNA variations and to understand their influence on strain functional divergence. RESULTS: There are 1,579 genes and 194 non-coding homologous sequences present in the genomes of all four strains, representing a 76. 2% conservation of the sequenced genome. About 60% of the X. fastidiosa unique sequences exist as tandem gene clusters of 6 or more genes. Multiple alignments identified 12,754 SNPs and 14,449 INDELs in the 1528 common genes and 20,779 SNPs and 10,075 INDELs in the 194 non-coding sequences. The average SNP frequency was 1.08 x 10(-2) per base pair of DNA and the average INDEL frequency was 2.06 x 10(-2) per base pair of DNA. On an average, 60.33% of the SNPs were synonymous type while 39.67% were non-synonymous type. The mutation frequency, primarily in the form of external INDELs was the main type of sequence variation. The relative similarity between the strains was discussed according to the INDEL and SNP differences. The number of genes unique to each strain were 60 (9a5c), 54 (Dixon), 83 (Ann1) and 9 (Temecula-1). A sub-set of the strain specific genes showed significant differences in terms of their codon usage and GC composition from the native genes suggesting their xenologous origin. Tandem repeat analysis of the genomic sequences of the four strains identified associations of repeat sequences with hypothetical and phage related functions. CONCLUSION: INDELs and strain specific genes have been identified as the main source of variations among strains, with individual strains showing different rates of genome evolution. Based on these genome comparisons, it appears that the Pierce's disease strain Temecula-1 genome represents the ancestral genome of the X. fastidiosa. Results of this analysis are publicly available in the form of a web database.

Analysis of Variance↗

Small molecules, big players: the National Cancer Institute's Initiative for Chemical Genetics.

In 2002, the National Cancer Institute created the Initiative for Chemical Genetics (ICG), to enable public research using small molecules to accelerate the discovery of cancer-relevant small-molecule probes. The ICG is a public-access research facility consisting of a tightly integrated team of synthetic and analytical chemists, assay developers, high-throughput screening and automation engineers, computational scientists, and software developers. The ICG seeks to facilitate the cross-fertilization of synthetic chemistry and cancer biology by creating a research environment in which new scientific collaborations are possible. To date, the ICG has interacted with 76 biology laboratories from 39 institutions and more than a dozen organic synthetic chemistry laboratories around the country and in Canada. All chemistry and screening data are deposited into the ChemBank web site (http://chembank.broad.harvard.edu/) and are available to the entire research community within a year of generation. ChemBank is both a data repository and a data analysis environment, facilitating the exploration of chemical and biological information across many different assays and small molecules. This report outlines how the ICG functions, how researchers can take advantage of its screening, chemistry and informatic capabilities, and provides a brief summary of some of the many important research findings.

Antineoplastic Agents↗

A new mixture model approach to analyzing allelic-loss data using Bayes factors.

BACKGROUND: Allelic-loss studies record data on the loss of genetic material in tumor tissue relative to normal tissue at various loci along the genome. As the deletion of a tumor suppressor gene can lead to tumor development, one objective of these studies is to determine which, if any, chromosome arms harbor tumor suppressor genes. RESULTS: We propose a large class of mixture models for describing the data, and we suggest using Bayes factors to select a reasonable model from the class in order to classify the chromosome arms. Bayes factors are especially useful in the case of testing that the number of components in a mixture model is n0 versus n1. In these cases, frequentist test statistics based on the likelihood ratio statistic have unknown distributions and are therefore not applicable. Our simulation study shows that Bayes factors favor the right model most of the time when tumor suppressor genes are present. When no tumor suppressor genes are present and background allelic-loss varies, the Bayes factors are often inconclusive, although this results in a markedly reduced false-positive rate compared to that of standard frequentist approaches. Application of our methods to three data sets of esophageal adenocarcinomas yields interesting differences from those results previously published. CONCLUSIONS: Our results indicate that Bayes factors are useful for analyzing allelic-loss data.

Adenocarcinoma↗

GANN: genetic algorithm neural networks for the detection of conserved combinations of features in DNA.

BACKGROUND: The multitude of motif detection algorithms developed to date have largely focused on the detection of patterns in primary sequence. Since sequence-dependent DNA structure and flexibility may also play a role in protein-DNA interactions, the simultaneous exploration of sequence- and structure-based hypotheses about the composition of binding sites and the ordering of features in a regulatory region should be considered as well. The consideration of structural features requires the development of new detection tools that can deal with data types other than primary sequence. RESULTS: GANN (available at http://bioinformatics.org.au/gann) is a machine learning tool for the detection of conserved features in DNA. The software suite contains programs to extract different regions of genomic DNA from flat files and convert these sequences to indices that reflect sequence and structural composition or the presence of specific protein binding sites. The machine learning component allows the classification of different types of sequences based on subsamples of these indices, and can identify the best combinations of indices and machine learning architecture for sequence discrimination. Another key feature of GANN is the replicated splitting of data into training and test sets, and the implementation of negative controls. In validation experiments, GANN successfully merged important sequence and structural features to yield good predictive models for synthetic and real regulatory regions. CONCLUSION: GANN is a flexible tool that can search through large sets of sequence and structural feature combinations to identify those that best characterize a set of sequences.

Algorithms↗