Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

[Modeling of all genome and database].

We have developed the protein modeling software FAMS (Full automatic protein modeling system), and using the FAMS the proteins coded in the all the genes were modeled. And we developed web browsing software. We had participated in the CAFASP2 contest of the CASP4 which is the competition of the protein structure prediction. We won almost best server in the CAFASP2 which is the contest of full automatic protein modeling. Accordingly the database quality made by using the FAMS program will be very good. The FAMS modeling web service is available in http://physchem.pharm.kitasato-u.ac.jp/. FAMSBASE is seen in the web site of http://famsbase.bio.nagoya-u.ac.jp/.

Databases, Genetic↗

From transcriptomics to bibliomics.

BACKGROUND: Current biological investigations tend to operate with genomes, instead of genes as during the last century. It is possible to compare entire genomes, transcriptomes or proteomes, using alphanumeric data corresponding to the differential expression levels of thousands of genes. What remains difficult is to link array results to factual or bibliographical data and retrieve information that is highly structured and - in Shannon's sense - rare. MATERIAL/METHODS: We have developed a tool, Documentation and Information LIBrary (DILIB), that enables us to retrieve, organize and analyze huge amounts of data available on the Internet and related to microarray experiments. DILIB can link hundreds of differentially expressed genes - through their Single Identifier or GenBank accession number - to hundreds of Medline records, and can retrieve, analyze, and compare automatically thousands of non-trivial descriptors related to gene clusters. RESULTS: As exemplified with frequency comparison of MEdical Subject Headings and Registry Number descriptors, we reanalyzed the involvement of 'integrin', 'interleukin' and 'CD Antigens' in mesotheliomas. Thus, DILIB allowed us to: (i). associate literature to expressed genes, (ii). link functional transcriptomes in various experiments, (iii). associate specific descriptors to experiments, (iv). define new research areas, and eventually (v). find new functions for co-expressed genes. CONCLUSIONS: We propose a new concept, 'bibliomics', representing a subset of high quality and rare information, retrieved and organized by systematic literature-searching tools from existing databases, and related to a subset of genes functioning together in '-omic' sciences.

Databases, Genetic↗

Annotating significant pairs of transcription factor binding sites in regulatory DNA.

In the presented work we search for transcription factor binding sites (BS) by including additional information about typical BS patterns. The new proposed score combines the ordinary profile score based on TRANSFAC-matrices together with a score based on pairs of BS. The latter score positively weights pairs of BS that tend to occur together in many regulatory DNA-sequences, in contrast to a random background model. The empirical BS pair frequencies result from our evaluation of a large dataset of orthologous genes.

Animals↗

[The study of HLA-Cw polymorphism in Uygur population].

The HLA-Cw loci polymorphism in Uygur population was investigated using the PCR- sequence specific oligonucleotide probe (SSOP) method,and the genetic database on the distribution of gene frequency of the HLA-Cw loci was established. From 146 individuals of Uygur population,18 HLA-Cw alleles were detected. The gene frequency was from 0.0069 to 0.2460. The four most common alleles were HLA-Cw*04(24.60%),07(11.51%),08(10 10%),14(12.02%),and they covered 58.23% of total alleles detected from Uygur population.We have made a survey of HLA-Cw alleles frequencies in a Uygur population,with blank frequency being lowered to 0.0064. The distribution of genotype frequencies met the law of Hardy-Weinberg equilibrium by hi-square test. The frequency data can be used in forensic and paternity tests to estimate the frequency of a DNA profile in the Uygur population,transplant matching and anthropology.

English Abstract↗

[The application of human mutation databases].

Researches on genome mutation are becoming more and more important with the finish of human genome DNA draft. This review is to classify the existing human mutation databases, including mutation database, SNP(single nucleotide polymorphisms) databases, mutation databases about disease, mutation databases about proteins, mutation databases about map and mutation information about specific gene. We also give advice on how to utilize these mutation databases, and discuss problems of existing databases.

Databases, Factual↗

Using protein motif combinations to update KEGG pathway maps and orthologue tables.

We have studied the projection of protein family data onto single bacterial translated genome as a solution to visualise relationships between families restricted to bacterial sequences. Any member of any type of family as defined in the Pfam database (domains, signatures, etc.) is considered as a protein module. Our first goal is to discover rules correlating the occurrence of modules with biochemical properties. To achieve this goal we have developed a platform to quantify information found in protein databases and to support the analysis of the nature of modules, their position and corresponding frequencies of occurrence (in isolation or in combination) in association with pathway knowledge as found in KEGG. This paper focuses on two pathways: the two-component system and the aminophosphonate metabolism, that are partially but not completely documented. Proteins involved in those pathways were listed separately in each organism to analyse module composition and rules constraining pathway interactions were identified. It is shown how these results can be used to update KEGG pathways and orthologue tables.

Animals↗

Linking experimental results, biological networks and sequence analysis methods using Ontologies and Generalised Data Structures.

The structure of a closely integrated data warehouse is described that is designed to link different types and varying numbers of biological networks, sequence analysis methods and experimental results such as those coming from microarrays. The data schema is inspired by a combination of graph based methods and generalised data structures and makes use of ontologies and meta-data. The core idea is to consider and store biological networks as graphs, and to use generalised data structures (GDS) for the storage of further relevant information. This is possible because many biological networks can be stored as graphs: protein interactions, signal transduction networks, metabolic pathways, gene regulatory networks etc. Nodes in biological graphs represent entities such as promoters, proteins, genes and transcripts whereas the edges of such graphs specify how the nodes are related. The semantics of the nodes and edges are defined using ontologies of node and relation types. Besides generic attributes that most biological entities possess (name, attribute description), further information is stored using generalised data structures. By directly linking to underlying sequences (exons, introns, promoters, amino acid sequences) in a systematic way, close interoperability to sequence analysis methods can be achieved. This approach allows us to store, query and update a wide variety of biological information in a way that is semantically compact without requiring changes at the database schema level when new kinds of biological information is added. We describe how this datawarehouse is being implemented by extending the text-mining framework ONDEX to link, support and complement different bioinformatics applications and research activities such as microarray analysis, sequence analysis and modelling/simulation of biological systems. The system is developed under the GPL license and can be downloaded from http://sourceforge.net/projects/ondex/

Algorithms↗

Deriving an ontology for human gene expression sources from the CYTOMER database on human organs and cell types.

CYTOMER is a relational database of organs/tissues, cell types, physiological systems and developmental stages that currently focuses on the human system. From this database, we have derived an ontology for anatomical and morphological structures for the human organism which includes all embryonal stages and the cell types constituting these structures. The ontology has been transferred to the OWL format and is freely available for download at http://cytomer/bioinf.med.uni-goettingen.de.

Animals↗

Clinical bioinformatics.

Clinical bioinformatics provides biological and medical information to allow for individualized healthcare. In this review, we describe the uses of clinical bioinformatics. After the analysis of the complete human genome sequences, clinical bioinformatics enables researchers to search online biological databases and use the biological information in their medical practices. The data obtained from using microarray is extremely complicated. In clinical bioinformatics, selecting appropriate software to analyze the microarray data for medical decision making is crucial. Proteomics strategy tools usually focus on similarity searches, structure prediction, and protein modeling. In clinical bioinformatics, the proteomic data only have meaning if they are integrated with clinical data. In pharmacogenomics, clinical bioinformatics includes elaborate studies of bioinformatics tools and various facets of proteomics related to drug target identification and clinical validation. Using clinical bioinformatics, researchers apply computational and high-throughput experimental techniques to cancer research and systems biology. Meanwhile, researchers of bioinformatics and medical information have incorporated clinical bioinformatics to improve health care, using biological and medical information. Using the high volume of biological information from clinical bioinformatics will contribute to changes in practice standards in the healthcare system. We believe that clinical bioinformatics provides benefits of improving healthcare, disease prevention and health maintenance as we move toward the era of personalized medicine.

Computational Biology↗

[Strategy for the protein identification of human proteome expression profile: selection of searching database].

Widely used method of protein identification for high-throughout proteome expression profile studies was database-dependent, so the selection of databases for the protein identification was very important. Despite the deficiency of available human protein databases, the complementarity of human proteins could be got mainly from human genome but not from the protein databases of other organisms. According to the comparison of the current protein databases from different aspects, IPI was recommended for the basic identification for the studies of human proteome expression profile, and other human protein or nucleic acid databases were needed for the complementary identification and novel protein mining.

Animals↗

Designing new methodologies for integrating biomedical information in clinical trials.

OBJECTIVES: To propose a modification to current methodologies for clinical trials, improving data collection and cost-efficiency. To describe a system to integrate distributed and heterogeneous medical and genetic databases for improving information access, retrieval and analysis of biomedical information. METHODS: Data for clinical trials can be collected from remote, distributed and heterogeneous data sources. In this distributed scenario, we propose an ontologybased approach, with two basic operations: mapping and unification. Mapping outputs the semantic model of a virtual repository with the information model of a specific database. Unification provides a single schema for two or more previously available virtual repositories. In both processes, domain ontologies can improve other traditional approaches. RESULTS: Private clinical databases and public genomic and disease databases (e.g., OMIM, Prosite and others) were integrated. We successfully tested the system using thirteen databases containing clinical and biological information and biomedical vocabularies. CONCLUSIONS: We present a domain-independent approach to biomedical database integration, used in this paper as a reference for the design of future models of clinico-genomic trials where information will be integrated, retrieved and analyzed. Such an approach to biomedical data integration has been one of the goals of the IST INFOBIOMED Network of Excellence in Biomedical Informatics, funded by the European Commission, and the new ACGT (Advanced Clinico-Genomic Trials on Cancer) project, where the authors will apply these methods to research experiments.

Clinical Trials as Topic↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

Mining Alzheimer disease relevant proteins from integrated protein interactome data.

Huge unrealized post-genome opportunities remain in the understanding of detailed molecular mechanisms for Alzheimer Disease (AD). In this work, we developed a computational method to rank-order AD-related proteins, based on an initial list of AD-related genes and public human protein interaction data. In this method, we first collected an initial seed list of 65 AD-related genes from the OMIM database and mapped them to 70 AD seed proteins. We then expanded the seed proteins to an enriched AD set of 765 proteins using protein interactions from the Online Predicated Human Interaction Database (OPHID). We showed that the expanded AD-related proteins form a highly connected and statistically significant protein interaction sub-network. We further analyzed the sub-network to develop an algorithm, which can be used to automatically score and rank-order each protein for its biological relevance to AD pathways(s). Our results show that functionally relevant AD proteins were consistently ranked at the top: among the top 20 of 765 expanded AD proteins, 19 proteins are confirmed to belong to the original 70 AD seed protein set. Our method represents a novel use of protein interaction network data for Alzheimer disease studies and may be generalized for other disease areas in the future.

Algorithms↗

VNTR polymorphism in the Buenos Aires, Argentina, metropolitan population.

VNTR loci provide a wealth of information for human genetic research, ranging from gene mapping to paternity testing and forensic identification. In this study we report the construction, validation, and analysis of the first local genetic database for VNTR markers for Argentina. A sample of the metropolitan population of Buenos Aires was typed by means of six VNTR systems. Allele frequencies and expected heterozygosity were calculated. The sample set was further tested for departures from Hardy-Weinberg equilibrium and power of exclusion. Allele frequency distributions are compatible with previously reported data on Caucasian populations, and no departures from Hardy-Weinberg equilibrium were detected.

Argentina↗

PRUFILE: a clinical and laboratory database for the genetics centre.

The growing complexity and volume of workload in a Clinical Genetics Centre can rapidly swamp the available clerical facilities. The multiuser database described gives facilities not only for administrative control and documentation but also for the production of data for clinical and scientific analysis. The close link between clinical and laboratory databases gives great versatility and easy expansion as new tests and disciplines are applied to clinical genetic problems.

Clinical Laboratory Information Systems↗

FINDbase: a relational database recording frequencies of genetic defects leading to inherited disorders worldwide.

Frequency of INherited Disorders database (FINDbase) (http://www.findbase.org) is a relational database, derived from the ETHNOS software, recording frequencies of causative mutations leading to inherited disorders worldwide. Database records include the population and ethnic group, the disorder name and the related gene, accompanied by links to any corresponding locus-specific mutation database, to the respective Online Mendelian Inheritance in Man entries and the mutation together with its frequency in that population. The initial information is derived from the published literature, locus-specific databases and genetic disease consortia. FINDbase offers a user-friendly query interface, providing instant access to the list and frequencies of the different mutations. Query outputs can be either in a table or graphical format, accompanied by reference(s) on the data source. Registered users from three different groups, namely administrator, national coordinator and curator, are responsible for database curation and/or data entry/correction online via a password-protected interface. Databaseaccess is free of charge and there are no registration requirements for data querying. FINDbase provides a simple, web-based system for population-based mutation data collection and retrieval and can serve not only as a valuable online tool for molecular genetic testing of inherited disorders but also as a non-profit model for sustainable database funding, in the form of a 'database-journal'.

Databases, Genetic↗