Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genetic databases”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

[The genetic polymorphism of 9 short tandem repeat loci in Yi ethnic group of Yunnan in China].

OBJECTIVE: To study the short teadem repeat(STR) genetics structure of a Chinese Yunnan Yi racial group. METHODS: Genetic distributions for nine STR loci were determined based on STR gene scan marked by fluorescence. RESULTS: Sixty-nine alleles and 164 kinds of genotypes were detected and identified from 84 unrelated Yi racial individuals. The corresponding gene and genotype frequencies were in 0.0060-0.5060 or 0.0119-0.4167 respectively. The expected and observed genotype frequencies of nine STR loci were in accordance with the Hardy-Weinberg equilibrium(P>0.05). The statistical analyses of nine STR loci showed that PIC was distributed in 0.5804-0.8777, H was in 0.6507-0.8002, DP was in 0.7976-0.9558, EPP was in 0.5207-0.8386, except TPOX and THO1 loci. CONCLUSION: Above research data enrich the Chinese genetic database, and play an important role in Chinese genetic study and in forensic application.

Asian People↗

The FlyBase database of the Drosophila Genome Projects and community literature.

The FlyBase Drosophila genetics database and the public interfaces of the Berkeley Drosophila Genome Project (BDGP) and European Drosophila Genome Project (EDGP) are in the process of integrating. At present, the data of these projects are available from independent, but hyperlinked, WWW sites (FlyBase URL, http://flybase. bio.indiana.edu/; BDGP URL, http://fruitfly.berkeley.edu/; EDGP URL, http://edgp.ebi.ac.uk/ ). Because of the considerable overlap of data classes between the contributions of the Drosophila genome projects and the Drosophila community, the new and enlarged FlyBase consortium views the implementation of a single integrated Drosophila genomics/genetics server as essential to the scientific community. This integration will occur in a stepwise fashion over the next 1-2 years. In this report, the salient features of the current databases and how to interrogate and navigate the extensive data sets are discussed.

Animals↗

[Genetic polymorphism of 8 STR loci on short arm of chromosome 3].

OBJECTIVE: To get the genotype and allele frequency distributions of 8 short tandem repeat (STR) loci on chromosome 3p (D3S1297, D3S1489, D3S1266, D3S1568, D3S1289, D3S1300, D3S1285 and D3S3681) in Chinese Han population in Hunan area. METHODS: Blood samples were collected from the random Han individuals in Hunan and the whole genomic DNA was extracted. STR loci were amplified by multiplex-PCR technique and genotyped by ABI 377 sequencer. RESULTS: Ninety-one alleles were detected, with frequencies ranging from 0.002 to 0.431, and these alleles constituted 312 genotypes. All the 8 loci met Hardy-Weinberg equilibrium. The statistical analysis of 8 STR loci showed the heterozygosity (H) >or= 0.729, the discrimination power (DP) >or= 0.725, the probabilities of paternity exclusion (PPE) >or= 0.596, and the polymorphic information content (PIC >or= 0.682). The result indicated that there was a significant difference between Han ethnic group and the white and the black. CONCLUSION: These results could serve as valuable data to enrich the Chinese genetic database and play an important role in Chinese population genetic and forensic medical application.

Adult↗

[Study of genetic polymorphism of 6 short tandem repeat loci in Nongqu Mongolian of inner Mongolia Autonomous Region in China].

OBJECTIVE: To get the genotype and allele frequency distribution of 6 short tandem repeat (STR) loci VWA, FGA, PENTAE, D6S1043, D2S1772, D7S3048 in NongQu Mongolia of China. METHODS: Two hundred and ninety-three unrelated individuals from Nongqu Mongolian were investigated. Polymerase chain reaction and polyacrylamide gel electrophoresis were used. RESULTS: Eighty alleles and 335 genotypes were detected, with frequencies ranging from 0.0017 to 0.2828. All the 6 loci met Hardy-Weinberg equilibrium. The statistical analysis of 6 STR loci showed the heterozygosity (H) >/= 0.7945, the discrimination power (DP) >/= 0.9160, the probability of paternity exclusion (PPE) >/= 0.5919, and the polymorphic information content (PIC) >/= 0.7617. CONCLUSION: These results could serve as valuable data to enrich the Chinese genetic database and play an important role in Chinese population genetic forensic medical application.

Alleles↗

[Genetic polymorphisms of 9 STR loci in Dongxiang ethnic group of China].

Genetic distribution for nine STR loci was determined in a Chinese Dongxing ethnic group based on STR genescan marked by fluorescence. Seventy-two alleles and 182 genotypes were observed in 94 unrelated Chinese Dongxiang individuals,with the corresponding gene frequency and genotype frequency being 0.0053-0.5825 and 0.0106-0.2660 respectively. The genotypes of nine STR loci were in accordance with the Hardy-Weinberg equilibrium (P>0.05). The statistical analysis of nine STR loci showed PIC (polymorphism information content, PIC) = or > 0.6378, H(heterozygosity, H) = or > 0.6500, DP (discrimination power, DP) = or > 0.8216, PPE (probabilities of paternity exculation, PPE) = or > 0.4903. The result indicated that there was a significant difference between Dongxiang ethnic group and the white and the black. There was no significant difference in Han nationality. These result filled the Dongxiang ethnic group-a specific group of Chinese into the genetic database and played an important role in Chinese population genetic study and forensic medicine application.

English Abstract↗

The maize genetics and genomics database. The community resource for access to diverse maize data.

The Maize Genetics and Genomics Database (MaizeGDB) serves the maize (Zea mays) research community by making a wealth of genetics and genomics data available through an intuitive Web-based interface. The goals of the MaizeGDB project are 3-fold: to provide a central repository for public maize information; to present the data through the MaizeGDB Web site in a way that recapitulates biological relationships; and to provide an array of computational tools that address biological questions in an easy-to-use manner at the site. In addition to these primary tasks, MaizeGDB team members also serve the community of maize geneticists by lending technical support for community activities, including the annual Maize Genetics Conference and various workshops, teaching researchers to use both the MaizeGDB Web site and Community Curation Tools, and engaging in collaboration with individual research groups to make their unique data types available through MaizeGDB.

Base Sequence↗

Creation and maintenance of Helix, a Web based database of medical genetics laboratories, to serve the needs of the genetics community.

Helix (healthlinks.washington.edu/helix) is a web accessible database that serves as the main U.S. directory of laboratories offering genetic testing. The database was designed to address the previously unmet need for a centralized, continuously updated source of information about clinical and research genetic testing to keep pace with the rapid rate of gene discovery resulting from the Human Genome Project. The Helix project began in 1992 at the University of Washington and Children's Hospital and Regional Medical Center. It has evolved from a single user stand alone relational database to a fully Web enabled database queried and maintained via the web and linked to other web accessible genomic databases. As of February, 1998 it lists more than 500 diseases and 290 laboratories, with over 5,200 registered users making approximately 250 queries/day (90% via the Internet). We describe the iterative design, implementation, population and assessment of the database over a six year period.

Database Management Systems↗

GenomEUtwin: a strategy to identify genetic influences on health and disease.

In this issue of Twin Research, we describe different facets of a European Community funded effort, GenomEUtwin, which capitalises on eight of the world's largest and best characterised twin registers and a multi-national population cohort, MORGAM. This international study, reaching beyond the geographical borders of Europe, is based on linkage and association strategies designed to identify genetic contributors to health and disease using integrated expertise of participating groups in genetics, epidemiology and biostatistics. By merging information from numerous epidemiological and genetic databases, GenomEUtwin will create an intellectual and technical infrastructure for future genetic epidemiological studies aiming to define genetic and life style risks for common human diseases.

European Union↗

Locus-specific databases: from ethical principles to practice.

Locus-specific databases (LSDBs) play an essential role in clinical care and research. They differ from traditional genetic databases in that they propose to place the mutations of "anonymized" patients directly on the World Wide Web. The proliferation of ethical guidelines and legal requirements affects the rapid and free transmission of clinical data, which is vital for both the daily management of patients and research into better diagnostics and treatment. This paper proposes a review of ethical principles endorsed by international instruments that are of particular relevance to LSDBs. It aims to translate them into 12 proposed practical guidelines that LSDB curators can use in collecting data for clinical research. Perhaps these guideposts will serve as a first step toward translating principles into practice.

Databases, Nucleic Acid↗

A measure of DNA sequence dissimilarity based on Mahalanobis distance between frequencies of words.

A number of algorithms exist for searching genetic databases for biologically significant similarities in DNA sequences. Past research has shown that word-based search tools are computationally efficient and can find similarities or dissimilarities invisible to other algorithms like FASTA. We characterize a family of word-based dissimilarity measures that define distance between two sequences by simultaneously comparing the frequencies of all subsequences of n adjacent letters (i.e., n-words) in the two sequences. Applications to real data demonstrate that currently used word-based methods that rely on Euclidean distance can be significantly improved by using Mahalanobis distance, which accounts for both variances and covariances between frequencies of n-words. Furthermore, in those cases where Mahalanobis distance may be too difficult to compute, using standardized Euclidean distance, which only corrects for the variances of frequencies of n-words, still gives better performance than the Euclidean distance. Also, a simple way of combining distances obtained at different n-words is considered. The goal is to obtain a single measure of dissimilarity between two DNA sequences. The performance ranking of the preceding three distances still holds for their combined counterparts. All results obtained in this paper are applicable to amino acid sequences with minor modifications.

Algorithms↗

The Genetic Activity Profile database.

A graphic approach termed a Genetic Activity Profile (GAP) has been developed to display a matrix of data on the genetic and related effects of selected chemical agents. The profiles provide a visual overview of the quantitative (doses) and qualitative (test results) data for each chemical. Either the lowest effective dose (LED) or highest ineffective dose (HID) is recorded for each agent and bioassay. Up to 200 different test systems are represented across the GAP. Bioassay systems are organized according to the phylogeny of the test organisms and the end points of genetic activity. The methodology for the production and evaluation of GAPs has been developed in collaboration with the International Agency for Research on Cancer. Data on individual chemicals have been compiled by IARC and by the U.S. Environmental Protection Agency. Data are available on 299 compounds selected from volumes 1-50 of the IARC Monographs and on 115 compounds identified as Superfund Priority Substances. Software to display the GAPs on an IBM-compatible personal computer is available from the authors. Structurally similar compounds frequently display qualitatively and quantitatively similar GAPs. By examining the patterns of GAPs of pairs and groups of chemicals, it is possible to make more informed decisions regarding the selection of test batteries to be used in evaluating chemical analogs. GAPs have provided useful data for the development of weight-of-evidence hazard ranking schemes. Also, some knowledge of the potential genetic activity of complex environmental mixtures may be gained from assessing the GAPs of component chemicals. The fundamental techniques and computer programs devised for the GAP database may be used to develop similar databases in other disciplines.

Animals↗

PAQ: Partition Analysis of Quasispecies.

MOTIVATION: The complexities of genetic data may not be accurately described by any single analytical tool. Phylogenetic analysis is often used to study the genetic relationship among different sequences. Evolutionary models and assumptions are invoked to reconstruct trees that describe the phylogenetic relationship among sequences. Genetic databases are rapidly accumulating large amounts of sequences. Newly acquired sequences, which have not yet been characterized, may require preliminary genetic exploration in order to build models describing the evolutionary relationship among sequences. There are clustering techniques that rely less on models of evolution, and thus may provide nice exploratory tools for identifying genetic similarities. Some of the more commonly used clustering methods perform better when data can be grouped into mutually exclusive groups. Genetic data from viral quasispecies, which consist of closely related variants that differ by small changes, however, may best be partitioned by overlapping groups. RESULTS: We have developed an intuitive exploratory program, Partition Analysis of Quasispecies (PAQ), which utilizes a non-hierarchical technique to partition sequences that are genetically similar. PAQ was used to analyze a data set of human immunodeficiency virus type 1 (HIV-1) envelope sequences isolated from different regions of the brain and another data set consisting of the equine infectious anemia virus (EIAV) regulatory gene rev. Analysis of the HIV-1 data set by PAQ was consistent with phylogenetic analysis of the same data, and the EIAV rev variants were partitioned into two overlapping groups. PAQ provides an additional tool which can be used to glean information from genetic data and can be used in conjunction with other tools to study genetic similarities and genetic evolution of viral quasispecies.

Algorithms↗

A founder mutation in presenilin 1 causing early-onset Alzheimer disease in unrelated Caribbean Hispanic families.

CONTEXT: Genetic determinants of Alzheimer disease (AD) have not been comprehensively examined in Caribbean Hispanics, a population in the United States in whom the frequency of AD is higher compared with non-Hispanic whites. OBJECTIVE: To identify variant alleles in genes related to familial early-onset AD among Caribbean Hispanics. DESIGN AND SETTING: Family-based case series conducted in 1998-2001 at an AD research center in New York, NY, and clinics in the Dominican Republic. PATIENTS: Among 206 Caribbean Hispanic families with 2 or more living members with AD who were identified, 19 (9.2%) had at least 1 individual with onset of AD before the age of 55 years. MAIN OUTCOME MEASURE: The entire coding region of the presenilin 1 gene and exons 16 and 17 of the amyloid precursor protein gene were sequenced in probands from the 19 families and their living relatives. RESULTS: A G-to-C nucleotide change resulting in a glycine-alanine amino acid substitution at codon 206 (Gly206Ala) in exon 7 of presenilin 1 was observed in 23 individuals from 8 (42%) of the 19 families. A Caribbean Hispanic individual with the Gly206Ala mutation and early-onset familial disease was also found by sequencing the corresponding genes of 319 unrelated individuals in New York City. The Gly206Ala mutation was not found in public genetic databases but was reported in 5 individuals from 4 Hispanic families with AD referred for genetic testing. None of the members of these families were related to one another, yet all carriers of the Gly206Ala mutation tested shared a variant allele at 2 nearby microsatellite polymorphisms, indicating a common ancestor. No mutations were found in the amyloid precursor protein gene. CONCLUSIONS: The Gly206Ala mutation was found in 8 of 19 unrelated Caribbean Hispanic families with early-onset familial AD. This genetic change may be a prevalent cause of early-onset familial AD in the Caribbean Hispanic population.

Age of Onset↗

HbVar: A relational database of human hemoglobin variants and thalassemia mutations at the globin gene server.

We have constructed a relational database of hemoglobin variants and thalassemia mutations, called HbVar, which can be accessed on the web at http://globin.cse.psu.edu. Extensive information is recorded for each variant and mutation, including a description of the variant and associated pathology, hematology, electrophoretic mobility, methods of isolation, stability information, ethnic occurrence, structure studies, functional studies, and references. The initial information was derived from books by Dr. Titus Huisman and colleagues [Huisman et al., 1996, 1997, 1998]. The current database is updated regularly with the addition of new data and corrections to previous data. Queries can be formulated based on fields in the database. Tables of common categories of variants, such as all those involving the alpha1-globin gene (HBA1) or all those that result in high oxygen affinity, are maintained by automated queries on the database. Users can formulate more precise queries, such as identifying "all beta-globin variants associated with instability and found in Scottish populations." This new database should be useful for clinical diagnosis as well as in fundamental studies of hemoglobin biochemistry, globin gene regulation, and human sequence variation at these loci.

Databases, Genetic↗

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗

Discovering disease-genes by topological features in human protein-protein interaction network.

MOTIVATION: Mining the hereditary disease-genes from human genome is one of the most important tasks in bioinformatics research. A variety of sequence features and functional similarities between known human hereditary disease-genes and those not known to be involved in disease have been systematically examined and efficient classifiers have been constructed based on the identified common patterns. The availability of human genome-wide protein-protein interactions (PPIs) provides us with new opportunity for discovering hereditary disease-genes by topological features in PPIs network. RESULTS: This analysis reveals that the hereditary disease-genes ascertained from OMIM in the literature-curated (LC) PPIs network are characterized by a larger degree, tendency to interact with other disease-genes, more common neighbors and quick communication to each other whereas those properties could not be detected from the network identified from high-throughput yeast two-hybrid mapping approach (EXP) and predicted interactions (PDT) PPIs network. KNN classifier based on those features was created and on average gained overall prediction accuracy of 0.76 in cross-validation test. Then the classifier was applied to 5262 genes on human genome and predicted 178 novel disease-genes. Some of the predictions have been validated by biological experiments.

Chromosome Mapping↗

Predicting deleterious nsSNPs: an analysis of sequence and structural attributes.

BACKGROUND: There has been an explosion in the number of single nucleotide polymorphisms (SNPs) within public databases. In this study we focused on non-synonymous protein coding single nucleotide polymorphisms (nsSNPs), some associated with disease and others which are thought to be neutral. We describe the distribution of both types of nsSNPs using structural and sequence based features and assess the relative value of these attributes as predictors of function using machine learning methods. We also address the common problem of balance within machine learning methods and show the effect of imbalance on nsSNP function prediction. We show that nsSNP function prediction can be significantly improved by 100% undersampling of the majority class. The learnt rules were then applied to make predictions of function on all nsSNPs within Ensembl. RESULTS: The measure of prediction success is greatly affected by the level of imbalance in the training dataset. We found the balanced dataset that included all attributes produced the best prediction. The performance as measured by the Matthews correlation coefficient (MCC) varied between 0.49 and 0.25 depending on the imbalance. As previously observed, the degree of sequence conservation at the nsSNP position is the single most useful attribute. In addition to conservation, structural predictions made using a balanced dataset can be of value. CONCLUSION: The predictions for all nsSNPs within Ensembl, based on a balanced dataset using all attributes, are available as a DAS annotation. Instructions for adding the track to Ensembl are at http://www.brightstudy.ac.uk/das_help.html.

Algorithms↗

GeneSeer: a sage for gene names and genomic resources.

BACKGROUND: Independent identification of genes in different organisms and assays has led to a multitude of names for each gene. This balkanization makes it difficult to use gene names to locate genomic resources, homologs in other species and relevant publications. METHODS: We solve the naming problem by collecting data from a variety of sources and building a name-translation database. We have also built a table of homologs across several model organisms: H. sapiens, M. musculus, R. norvegicus, D. melanogaster, C. elegans, S. cerevisiae, S. pombe and A. thaliana. This allows GeneSeer to draw phylogenetic trees and identify the closest homologs. This, in turn, allows the use of names from one species to identify homologous genes in another species. A website http://geneseer.cshl.org/ is connected to the database to allow user-friendly access to our tools and external genomic resources using familiar gene names. CONCLUSION: GeneSeer allows access to gene information through common names and can map sequences to names. GeneSeer also allows identification of homologs and paralogs for a given gene. A variety of genomic data such as sequences, SNPs, splice variants, expression patterns and others can be accessed through the GeneSeer interface. It is freely available over the web http://geneseer.cshl.org/ and can be incorporated in other tools through an http-based software interface described on the website. It is currently used as the search engine in the RNAi codex resource, which is a portal for short hairpin RNA (shRNA) gene-silencing constructs.

Alternative Splicing↗