Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Large-scale validation of single nucleotide polymorphisms in gene regions.

Genome-wide association studies using large numbers of bi-allelic single nucleotide polymorphisms (SNPs) have been proposed as a potentially powerful method for identifying genes involved in common diseases. To assemble a SNP collection appropriate for large-scale association, we designed assays for 226,099 publicly available SNPs located primarily within known and predicted gene regions. Allele frequencies were estimated in a sample of 92 CEPH Caucasians using chip-based MALDI-TOF mass spectrometry with pooled DNA. Of the 204,200 designed assays that were functional, 125,799 SNPs were determined to be polymorphic (minor allele frequency > 0.02), of which 101,729 map uniquely to the human genome. Many of the commonly available RefSNP annotations were predictive of polymorphic status and could be used to improve the selection of SNPs from the public domain for genetic research. The set of uniquely mapping, polymorphic SNPs is located within 10 kb of 66% of known and predicted genes annotated in LocusLink, which could prove useful for large-scale disease association studies.

Databases, Genetic↗

A new approach to decoding life: systems biology.

Systems biology studies biological systems by systematically perturbing them (biologically, genetically, or chemically); monitoring the gene, protein, and informational pathway responses; integrating these data; and ultimately, formulating mathematical models that describe the structure of the system and its response to individual perturbations. The emergence of systems biology is described, as are several examples of specific systems approaches.

Animals↗

Intrachromosomal serial replication slippage in trans gives rise to diverse genomic rearrangements involving inversions.

Serial replication slippage in cis (SRScis) provides a plausible explanation for many complex genomic rearrangements that underlie human genetic disease. This concept, taken together with the intra- and intermolecular strand switch models that account for mutations that arise via quasipalindrome correction, suggest that intrachromosomal SRS in trans (SRStrans) mediated by short inverted repeats may also give rise to a diverse series of complex genomic rearrangements. If this were to be so, such rearrangements would invariably generate inversions. To test this idea, we collated all informative mutations involving inversions of >or=5 bp but <1 kb by screening the Human Gene Mutation Database (HGMD; www.hgmd.org) and conducting an extensive literature search. Of the 21 resulting mutations, only two (both of which coincidentally contain untemplated additions) were found to be incompatible with the SRStrans model. Eighteen (one simple inversion, six inversions involving sequence replacement by upstream or downstream sequence, five inversions involving the partial reinsertion of removed sequence, and six inversions that occurred in a more complicated context) of the remaining 19 mutations were found to be consistent with either two steps of intrachromosomal SRStrans or a combination of replication slippage in cis plus intrachromosomal SRStrans. The remaining lesion, a 31-kb segmental duplication associated with a small inversion in the SLC3A1 gene, is explicable in terms of a modified SRS model that integrates the concept of "break-induced replication." This study therefore lends broad support to our postulate that intrachromosomal SRStrans can account for a variety of complex gene rearrangements that involve inversions.

Amino Acid Transport Systems, Basic↗

There's no place like WormBase: an indispensable resource for Caenorhabditis elegans researchers.

The nematode Caenorhabditis elegans is used extensively by scientists to study a wide variety of biological processes and is one of the most thoroughly characterized animals. Over the years, the community of C. elegans researchers has generated a wealth of information on the genetics, development, behaviour, and cellular and molecular biology of the worm. This body of data has grown even larger with the recent application of high throughput screening methodology to study gene function, expression and interactions. WormBase (http://www.wormbase.org) is the primary online source of biological data on C. elegans and related nematodes. Equipped with an assortment of powerful search tools, WormBase allows users to quickly extract a variety of information, including data on individual genes, DNA sequence, cell lineage and literature citations. As the database is well maintained and the functionalities constantly modified in response to evolving researcher needs, WormBase has become a vital component of the laboratories studying the worm and a model for other biological databases.

Animals↗

Transcriptional regulation of nanog by OCT4 and SOX2.

Nanog, Sox2, and Oct4 are transcription factors all essential to maintaining the pluripotent embryonic stem cell phenotype. Through a cooperative interaction, Sox2 and Oct4 have previously been described to drive pluripotent-specific expression of a number of genes. We now extend the list of Sox2-Oct4 target genes to include Nanog. Within the Nanog proximal promoter, we identify a composite sox-oct cis-regulatory element essential for Nanog pluripotent transcription. This element is conserved over 250 million years of cumulative evolution within the eutherian mammals. A Nanog proximal promoter-EGFP (enhanced green fluorescent protein) reporter transgene recapitulates endogenous Nanog mRNA expression in embryonic stem cells and their differentiated derivatives. Sox2 and Oct4 interaction with the Nanog promoter was confirmed through mutagenesis and in vitro binding assays. Electrophoretic mobility shift assays indicate that the Sox2-Oct4 heterodimer forms more efficiently on the composite element within Nanog than the similar element within Fgf4. Using chromatin immunoprecipitation, we show that Oct4 and Sox2 bind to the Nanog promoter in living mouse and human embryonic stem cells. Furthermore, by specific knockdown of Oct4 and Sox2 mRNA by RNA interference in embryonic stem cells, we provide genetic evidence for a link between Oct4, Sox2, and the Nanog promoter. These studies extend the understanding of the pluripotent genetic regulatory network within which the Sox2-Oct4 complex are at the top of the regulatory hierarchy.

Animals↗

The LRC haplotype project: a resource for killer immunoglobulin-like receptor-linked association studies.

There is increasing evidence for epistatic interactions between gene products (e.g. KIR) encoded within the Leukocyte Receptor Complex (LRC) with those (e.g. HLA) of the Major Histocompatibility Complex (MHC), resulting in susceptibility to disease. Identification of such associations at the DNA level requires comprehensive knowledge of the genetic variation and haplotype structure of the underlying loci. The LRC haplotype project aims to provide this knowledge by sequencing common LRC haplotypes.

Chromosome Mapping↗

Selecting informative genes with parallel genetic algorithms in tissue classification.

Recent advances in biotechnology offer the ability to measure the levels of expression of thousands of genes in parallel. Analysis of such data can provide understanding and insight into gene function and regulatory mechanisms. Several machine learning approaches have been used to aid to understand the functions of genes. However, these tasks are made more difficult due to the noisy nature of array data and the overwhelming number of gene features. In this paper, we use the parallel genetic algorithm to filter out the informative genes relative to classification. By combing with the classification method proposed by Golub et al. and Slonim et al., we classify the data sets with tissues of different classes, and the preliminary results are presented in this paper.

Algorithms↗

Quantitative structure activity relationship of benzoxazinone derivatives as neuropeptide Y Y5 receptor antagonists.

Quantitative structure activity relationship (QSAR) has been established for 30 benzoxazinone derivatives acting as neuropeptide Y Y5 receptor antagonists. The genetic algorithm and multiple linear regression were used to generate the relationship between biological activity and calculated descriptors. Model with good statistical qualities was developed using four descriptors from topological, thermodynamic, spatial and electrotopological class. The validation of the model was done by cross validation, randomization and external test set prediction.

Algorithms↗

High HIV-1 genetic diversity in Cuba.

BACKGROUND: HIV-1 subtype B is largely predominant in the Caribbean, although other subtypes have been recently identified in Cuba. OBJECTIVES: To examine HIV-1 genetic diversity in Cuba. METHODS: The study enrolled 105 HIV-1-infected individuals, 93 of whom had acquired the infection in Cuba. DNA from peripheral blood mononuclear cells was used for polymerase chain reaction amplification and sequencing of pol (protease-reverse transcriptase) and env (V3 region) segments. Phylogenetic trees were constructed using the neighbour-joining method. Intersubtype recombination was analysed by bootscanning. RESULTS: Of the samples, 50 (48%) were of subtype B and 55 (52%) of diverse non-B subtypes and recombinant forms. Among non-B viruses, 12 were non-recombinant, belonging to six subtypes (C, D, F1, G, H and J), the most frequent of which was subtype G (n = 5). The remaining 43 (78%) non-B viruses were recombinant, with 14 different forms, the two most common of which were Dpol/Aenv (n = 21) and U(unknown)pol/Henv (n = 7), which grouped in respective monophyletic clusters. Twelve recombinant viruses were mosaics of different genetic forms circulating in Cuba. Overall, 21 genetic forms were identified, with all known HIV-1 group M subtypes present in Cuba, either as non-recombinant viruses or as segments of recombinant forms. Non-B subtype viruses were predominant among heterosexuals (72%) and B subtype viruses among homo- or bisexuals (63%). CONCLUSION: An extraordinarily high diversity of HIV-1 genetic forms, unparalleled in the Americas and comparable to that found in Central Africa, is present in Cuba.

Cuba↗

The Rice PIPELINE: a unification tool for plant functional genomics.

The Rice Genome Research Project in Japan performs genome sequencing and comprehensive expression profiling, constructs genetic and physical maps, collects full-length cDNAs and generates mutant lines, all aimed at improving the breeding of the rice plant as a food source. The National Institute of Agrobiological Sciences in Tsukuba, Japan, has accumulated numerous rice biological resources and has already successfully produced a high-quality genome sequence, a high-density genetic map with 3000 markers, 30,000 full-length cDNAs, over 700 expression profiles with a 9000 cDNA microarray and 15,000 flanking sequences with Tos17 insertions in about 3765 mutant lines from about 50,000 transposon insertion lines. These resources are available in the public domain. A new unification tool for functional genomics, called Rice PIPELINE, has also been developed for the dynamic collection and compilation of genomics data (genome sequences, full-length cDNAs, gene expression profiles, mutant lines, cis elements) from various databases. The mission of Rice PIPELINE is to provide a unique scientific resource that pools publicly available rice genomic data for search by clone sequence, clone name, GenBank accession number, or keyword. The web-based form of Rice PIPELINE is available at http://cdna01.dna.affrc.go.jp/PIPE/.

Computational Biology↗

Analysis of genetic variation in the GenomEUtwin project.

Multiallelic short tandem repeat polymorphisms, or microsatellites, are useful markers in genome wide scans to identify chromosomal regions containing genes underlying disease loci. The biallelic single nucleotide polymorphism (SNP) can be used to fine map previously identified large candidate regions or to test functional candidate genes by association analysis. In the GenomEUtwin project the population based impact of susceptibility genes for six multifactorial traits will be studied. A genome wide panel of informative human microsatellite markers will be analyzed by fluorescent capillary electrophoresis in well characterized twin and population samples. Contrary to microsatellites, selection of the most informative panels of SNPs is hampered by imperfect data on the allele frequencies and population distribution of SNPs markers in the databases. Therefore, selection of SNPs requires a substantial amount of bioinformatics, and, the SNPs need to be validated experimentally in the relevant populations prior to genotyping large sample sets. In the GenomEUtwin project, large scale genotyping of SNPs will be performed using the SNPstreamUHT and MassARRAY genotyping systems that are based on the primer extension reaction principle combined with fluorescent and mass spectrometric detection, respectively. Production of the genotyping data will be a joint effort by GenomEUtwin partners at the University of Helsinki, the National Public Health Institute in Helsinki, Finland and Uppsala University, Sweden. All genotyping data will be stored in a common database established specifically for the GenomEUtwin project, from where it can be accessed by the twin research centres that provided the samples for genotyping.

Databases, Genetic↗

The hidden value of missing genotypes.

Robotic systems allow vast genetic data sets to be generated automatically with little manual input, raising questions about accuracy. To test whether errors occur randomly, I used the frequencies of missing genotypes in a large human data set to construct a population tree. Remarkably, the gaps appear to carry as strong a phylogenetic signal as the actual data themselves.

Data Collection↗

Periodicity of DNA in exons.

BACKGROUND: The periodic pattern of DNA in exons is a known phenomenon. It was suggested that one of the initial causes of periodicity could be the universal (RNY)npattern (R = A or G, Y = C or U, N = any base) of ancient RNA. Two major questions were addressed in this paper. Firstly, the cause of DNA periodicity, which was investigated by comparisons between real and simulated coding sequences. Secondly, quantification of DNA periodicity was made using an evolutionary algorithm, which was not previously used for such purposes. RESULTS: We have shown that simulated coding sequences, which were composed using codon usage frequencies only, demonstrate DNA periodicity very similar to the observed in real exons. It was also found that DNA periodicity disappears in the simulated sequences, when the frequencies of codons become equal. Frequencies of the nucleotides (and the dinucleotide AG) at each location along phase 0 exons were calculated for C. elegans, D. melanogaster and H. sapiens. Two models were used to fit these data, with the key objective of describing periodicity. Both of the models showed that the best-fit curves closely matched the actual data points. The first dynamic period determination model consistently generated a value, which was very close to the period equal to 3 nucleotides. The second fixed period model, as expected, kept the period exactly equal to 3 and did not detract from its goodness of fit. CONCLUSIONS: Conclusion can be drawn that DNA periodicity in exons is determined by codon usage frequencies. It is essential to differentiate between DNA periodicity itself, and the length of the period equal to 3. Periodicity itself is a result of certain combinations of codons with different frequencies typical for a species. The length of period equal to 3, instead, is caused by the triplet nature of genetic code. The models and evolutionary algorithm used for characterising DNA periodicity are proven to be an effective tool for describing the periodicity pattern in a species, when a number of exons in the same phase are analysed.

Algorithms↗

Continued colonization of the human genome by mitochondrial DNA.

Integration of mitochondrial DNA fragments into nuclear chromosomes (giving rise to nuclear DNA sequences of mitochondrial origin, or NUMTs) is an ongoing process that shapes nuclear genomes. In yeast this process depends on double-strand-break repair. Since NUMTs lack amplification and specific integration mechanisms, they represent the prototype of exogenous insertions in the nucleus. From sequence analysis of the genome of Homo sapiens, followed by sampling humans from different ethnic backgrounds, and chimpanzees, we have identified 27 NUMTs that are specific to humans and must have colonized human chromosomes in the last 4-6 million years. Thus, we measured the fixation rate of NUMTs in the human genome. Six such NUMTs show insertion polymorphism and provide a useful set of DNA markers for human population genetics. We also found that during recent human evolution, Chromosomes 18 and Y have been more susceptible to colonization by NUMTs. Surprisingly, 23 out of 27 human-specific NUMTs are inserted in known or predicted genes, mainly in introns. Some individuals carry a NUMT insertion in a tumor-suppressor gene and in a putative angiogenesis inhibitor. Therefore in humans, but not in yeast, NUMT integrations preferentially target coding or regulatory sequences. This is indeed the case for novel insertions associated with human diseases and those driven by environmental insults. We thus propose a mutagenic phenomenon that may be responsible for a variety of genetic diseases in humans and suggest that genetic or environmental factors that increase the frequency of chromosome breaks provide the impetus for the continued colonization of the human genome by mitochondrial DNA.

Algorithms↗

The mouse as a model for human biology: a resource guide for complex trait analysis.

The mouse has been a powerful force in elucidating the genetic basis of human physiology and pathophysiology. From its beginnings as the model organism for cancer research and transplantation biology to the present, when dissection of the genetic basis of complex disease is at the forefront of genomics research, an enormous and remarkable mouse resource infrastructure has accumulated. This review summarizes those resources and provides practical guidelines for their use, particularly in the analysis of quantitative traits.

Animals↗

Human limb malformations; an approach to the molecular basis of development.

Analysis of human inherited limb malformations and of mouse mutants copying individual human mutations team up to promote the understanding of vertebrate limb development as a model for molecular regulatory interactions in animals. The strength of the human genetic contribution lies in the increasingly complete information on the human genome, transcriptome and proteome, as well as in the wealth of individual mutations interfering with limb development available for study. Based on the strong fundament of the human genome project, mapping and identification of novel genes associated with limb defects extends considerably the range of candidates beyond the repertoire of developmental genes and pathways known from animals. Attempts to correlate genotype and phenotype uncover a very broad range of genetic heterogeneity, i.e. different genes underlying the same phenotype, or allelic heterogeneity between families, i.e. clinically distinct phenotypes associated with mutations affecting the same gene. Mechanisms other than simple Mendelian inheritance have to be taken into consideration. Phenotypic variability within families might be explained by different modifying genes or environmental influence, whereas asymmetry of limb defects within one patient may be caused by epigenetic factors, such as somatic mosaicism or X-inactivation, or by non-genetic factors. The intimate knowledge of the genes and events governing limb pattern formation in humans and animals will elucidate the regulatory interactions underlying normal and pathological development, homeostasis, and repair, and thus propose targets for preventive measures and novel approaches to therapeutic intervention in the new era of molecular medicine.

Bardet-Biedl Syndrome↗

Genetic variability of hepatitis B virus isolates in Poland.

There is very limited knowledge about the genetic variability of HBV strains circulating in the population of Polish chronically infected HBV patients. The aim of this study was to analyse the phylogenetic relatedness and polymorphism in some functional domains of HBV genome among chronically infected patients from northern Poland. Fifty-one serum samples were included to analysis of HBV genomes due to the viral load sufficient for DNA preparation and sequencing. The sequences of the rt polymerase/S and preC/BCP regions of those isolates were analysed, compared to genome sequences of different variants of HBV from GenBank database and genetic relatedness of Polish genotypes to known reference strains was estimated. A phylogenetic tree of 41 analysed genotype A isolates as well as 8 genotype D strains was constructed showing relationship to know reference strains. Two isolates, initially classified as genotype F turned to be related to genotype H, newly described genotype deriving from genotype F, a very rare genotype in Europe. HBV genotypes' distribution pattern in Poland and phylogenetic relatedness seems to be different from our Eastern neighbours. Due to the fact that Poland is still ethnically uniform country, it is interesting to explore molecular epidemiology of HBV infections in our population.

Adult↗

Differential gene expression profiling in genetic and multifactorial cardiovascular diseases.

Gene expression profiling by microarray technologies has been successfully applied to study the transcriptional changes that occur in tissues such as heart, vessels and blood cells in different cardiovascular disorders. Such studies have been performed in human cardiovascular syndromes and in animal models with the aim of unraveling the complex molecular pictures underlying human pathophysiology. As already observed in cancer research, gene expression studies in humans may provide a finer molecular classification of patients with cardiovascular diseases and indicate new markers useful for prognostic and therapeutic strategies. In this paper, we present the findings obtained with microarray platforms to explore transcriptome alterations in cardiovascular diseases. To describe the potential of global expression profiling approach in this field, we have chosen to review the genomic findings obtained in some classic heart diseases with genetic transmission such as hyperthrophic cardiomyopathy and Fabry disease, together with findings obtained in common multifactorial cardiovascular disorders such as heart failure, atherosclerosis and infarction. Wherever feasible, we present the results obtained in patients together with those obtained in the corresponding animal and cellular models.

Animals↗