Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Population genetic data on loci LDLR, GYPA, HBGG, D7S8 and GC in the Bangkok population compared with rural Thais from Trat province.

Prior to the introduction of any DNA marker as a tool for person identification and paternity test in certain ethnic groups, a population genetic database should be constructed. Using multiplex primers in single tube polymerase chain amplification, 5 loci of unrelated genes in the PM Amplitype kit (Perkin Elmer) were studied in two Thai population groups: 228 DNA samples were extracted from blood collected at the Borai rural area in Trat province; another 123 DNA samples were collected at the outpatient clinic, Department of Forensic Medicine, King Chulalongkorn Memorial Hospital, Bangkok. Analysis of alleles and genotypes was performed after reversed dot blot hybridization of PCR products to allelic sequence specific probes immobilized on the membrane strip followed by nonradioisotopic detection according to the manufacturer's protocol. Population genetic statistic parameters including discrimination power (DP), the probability of matching (PM), power of exclusion for trio (PE trio) and typical paternity index (PI typical) were computed. Both Thai population groups showed no significant deviation from the Hardy Weinberg Expectation (HWE). The combined DP of all 5 loci in the PM Amplitype markers was 0.993636 for rural Thais and 0.994409 for Thais from Bangkok. The combined PM for rural Thais and those living in Bangkok was 0.006364 and 0.005591, respectively. The combined PE trio was 0.696825 and 0.698875 in both Thai population groups and the combined PI typical values were < 1.0. In conclusion, person identification using PM Amplitype DNA markers was efficient and satisfactory within certain limits. Hence, the application of PM Amplitype DNA markers for paternity tests should be cautiously considered and applied in combination with other parameters.

Alleles↗

Computational verification of protein-protein interactions by orthologous co-expression.

BACKGROUND: High-throughput methods identify an overwhelming number of protein-protein interactions. However, the limited accuracy of these methods results in the false identification of many spurious interactions. Accordingly, the resulting interactions are regarded as hypothetical and computational methods are needed to increase their confidence. Several methods have recently been suggested for this purpose including co-expression as a confidence measure for interacting proteins, but their performance is still quite poor. RESULTS: We introduce a novel computational method for verification of protein-protein interactions based on the co-expression of orthologs of interacting partners. The performance of our method is analysed using known S. cerevisiae interactions, and is shown to overcome limitations of previous methods. We present specific examples of known and putative interactions that are detected by our method and not by previous methods, and suggest that they represent transient interactions that might have been conserved and stabilized in other species. CONCLUSION: Co-expression of orthologous protein-pairs can be used to increase the confidence of hypothetical protein-protein interactions in S. cerevisiae as well as in other species. This approach may be especially useful for species with no available expression profiles and for transient interactions.

Animals↗

Evolutionary sequence analysis of complete eukaryote genomes.

BACKGROUND: Gene duplication and gene loss during the evolution of eukaryotes have hindered attempts to estimate phylogenies and divergence times of species. Although current methods that identify clusters of orthologous genes in complete genomes have helped to investigate gene function and gene content, they have not been optimized for evolutionary sequence analyses requiring strict orthology and complete gene matrices. Here we adopt a relatively simple and fast genome comparison approach designed to assemble orthologs for evolutionary analysis. Our approach identifies single-copy genes representing only species divergences (panorthologs) in order to minimize potential errors caused by gene duplication. We apply this approach to complete sets of proteins from published eukaryote genomes specifically for phylogeny and time estimation. RESULTS: Despite the conservative criterion used, 753 panorthologs (proteins) were identified for evolutionary analysis with four genomes, resulting in a single alignment of 287,000 amino acids. With this data set, we estimate that the divergence between deuterostomes and arthropods took place in the Precambrian, approximately 400 million years before the first appearance of animals in the fossil record. Additional analyses were performed with seven, 12, and 15 eukaryote genomes resulting in similar divergence time estimates and phylogenies. CONCLUSION: Our results with available eukaryote genomes agree with previous results using conventional methods of sequence data assembly from genomes. They show that large sequence data sets can be generated relatively quickly and efficiently for evolutionary analyses of complete genomes.

Animals↗

Integrative missing value estimation for microarray data.

BACKGROUND: Missing value estimation is an important preprocessing step in microarray analysis. Although several methods have been developed to solve this problem, their performance is unsatisfactory for datasets with high rates of missing data, high measurement noise, or limited numbers of samples. In fact, more than 80% of the time-series datasets in Stanford Microarray Database contain less than eight samples. RESULTS: We present the integrative Missing Value Estimation method (iMISS) by incorporating information from multiple reference microarray datasets to improve missing value estimation. For each gene with missing data, we derive a consistent neighbor-gene list by taking reference data sets into consideration. To determine whether the given reference data sets are sufficiently informative for integration, we use a submatrix imputation approach. Our experiments showed that iMISS can significantly and consistently improve the accuracy of the state-of-the-art Local Least Square (LLS) imputation algorithm by up to 15% improvement in our benchmark tests. CONCLUSION: We demonstrated that the order-statistics-based integrative imputation algorithms can achieve significant improvements over the state-of-the-art missing value estimation approaches such as LLS and is especially good for imputing microarray datasets with a limited number of samples, high rates of missing data, or very noisy measurements. With the rapid accumulation of microarray datasets, the performance of our approach can be further improved by incorporating larger and more appropriate reference datasets.

Algorithms↗

[Genetic polymorphism of 9 STR loci in Chaoxian National Minority of China].

In order to enrich the Chinese genetic database,nine polymorphic loci of STR,such as D3S1358,vWA,FGA,TH01,TPOX,CSF1PO,D5S818,D13S317 and D7S820 were studied. Based on STR gene scan marked by fluorescence,91 unrelated Chinese Chaoxian individuals were observed.81 alleles and 196 genotypes were found. The corresponding gene frequency and genotype frequency were 0.0055-0.4615 and 0.0110-0.9890 respectively. The genogypes frequency of nine STR loci was good with the Hardy-Weinberg equilibrium (P approximately 0.05). The statistical analysis of nine STR loci showed the following: PIC (polymorphic information content) >or=0.6863, H (heterozygosity) >or=0.6919, DP (discrimination power) >or=0.8301, EPP (probability of paternity exclusion) >or=0.8590. The data studied can be used in Chinese population genetic studies and forensic medicine applications.

English Abstract↗

A statistical framework for genomic data fusion.

MOTIVATION: During the past decade, the new focus on genomics has highlighted a particular challenge: to integrate the different views of the genome that are provided by various types of experimental data. RESULTS: This paper describes a computational framework for integrating and drawing inferences from a collection of genome-wide measurements. Each dataset is represented via a kernel function, which defines generalized similarity relationships between pairs of entities, such as genes or proteins. The kernel representation is both flexible and efficient, and can be applied to many different types of data. Furthermore, kernel functions derived from different types of data can be combined in a straightforward fashion. Recent advances in the theory of kernel methods have provided efficient algorithms to perform such combinations in a way that minimizes a statistical loss function. These methods exploit semidefinite programming techniques to reduce the problem of finding optimizing kernel combinations to a convex optimization problem. Computational experiments performed using yeast genome-wide datasets, including amino acid sequences, hydropathy profiles, gene expression data and known protein-protein interactions, demonstrate the utility of this approach. A statistical learning algorithm trained from all of these data to recognize particular classes of proteins--membrane proteins and ribosomal proteins--performs significantly better than the same algorithm trained on any single type of data. AVAILABILITY: Supplementary data at http://noble.gs.washington.edu/proj/sdp-svm

Algorithms↗

Virtual Footprint and PRODORIC: an integrative framework for regulon prediction in prokaryotes.

SUMMARY: A new online framework for the accurate and integrative prediction of transcription factor binding sites (TFBSs) in prokaryotes was developed. The system consists of three interconnected modules: (1) The PRODORIC database as a comprehensive data source and extensive collection of TFBSs with corresponding position weight matrices. (2) The pattern matching tool Virtual Footprint for the prediction of genome based regulons and for the analysis of individual promoter regions. (3) The interactive genome browser GBPro for the visualization of TFBS search results in their genomic context and links to gene and regulator-specific information in PRODORIC. The aim of this service is to provide researchers a free and easy to use collection of interconnected tools in the field of molecular microbiology, infection and systems biology. AVAILABILITY: http://www.prodoric.de/vfp.

Algorithms↗

GeneNotes--a novel information management software for biologists.

BACKGROUND: Collecting and managing information is a challenging task in a genome-wide profiling research project. Most databases and online computational tools require a direct human involvement. Information and computational results are presented in various multimedia formats (e.g., text, image, PDF, word files, etc.), many of which cannot be automatically processed by computers in biologically meaningful ways. In addition, the quality of computational results is far from perfect and requires nontrivial manual examination. The timely selection, integration and interpretation of heterogeneous biological information still heavily rely on the sensibility of biologists. Biologists often feel overwhelmed by the huge amount of and the great diversity of distributed heterogeneous biological information. DESCRIPTION: We developed an information management application called GeneNotes. GeneNotes is the first application that allows users to collect and manage multimedia biological information about genes/ESTs. GeneNotes provides an integrated environment for users to surf the Internet, collect notes for genes/ESTs, and retrieve notes. GeneNotes is supported by a server that integrates gene annotations from many major databases (e.g., HGNC, MGI, etc.). GeneNotes uses the integrated gene annotations to (a) identify genes given various types of gene IDs (e.g., RefSeq ID, GenBank ID, etc.), and (b) provide quick views of genes. GeneNotes is free for academic usage. The program and the tutorials are available at: http://bayes.fas.harvard.edu/genenotes/. CONCLUSIONS: GeneNotes provides a novel human-computer interface to assist researchers to collect and manage biological information. It also provides a platform for studying how users behave when they manipulate biological information. The results of such study can lead to innovation of more intelligent human-computer interfaces that greatly shorten the cycle of biology research.

Biology↗

Towards precise classification of cancers based on robust gene functional expression profiles.

BACKGROUND: Development of robust and efficient methods for analyzing and interpreting high dimension gene expression profiles continues to be a focus in computational biology. The accumulated experiment evidence supports the assumption that genes express and perform their functions in modular fashions in cells. Therefore, there is an open space for development of the timely and relevant computational algorithms that use robust functional expression profiles towards precise classification of complex human diseases at the modular level. RESULTS: Inspired by the insight that genes act as a module to carry out a highly integrated cellular function, we thus define a low dimension functional expression profile for data reduction. After annotating each individual gene to functional categories defined in a proper gene function classification system such as Gene Ontology applied in this study, we identify those functional categories enriched with differentially expressed genes. For each functional category or functional module, we compute a summary measure (s) for the raw expression values of the annotated genes to capture the overall activity level of the module. In this way, we can treat the gene expressions within a functional module as an integrative data point to replace the multiple values of individual genes. We compare the classification performance of decision trees based on functional expression profiles with the conventional gene expression profiles using four publicly available datasets, which indicates that precise classification of tumour types and improved interpretation can be achieved with the reduced functional expression profiles. CONCLUSION: This modular approach is demonstrated to be a powerful alternative approach to analyzing high dimension microarray data and is robust to high measurement noise and intrinsic biological variance inherent in microarray data. Furthermore, efficient integration with current biological knowledge has facilitated the interpretation of the underlying molecular mechanisms for complex human diseases at the modular level.

Algorithms↗

[Genetic polymorphisms of 5 X-STR loci in Yunnan Nu population].

OBJECTIVE: To investigate the allele and genotype frequencies of DXS6804, DXS6799, DXS8378, DXS7130 and DXS7132 in unrelated individuals of Nu population and establish the related genetic database. METHODS: Five X-STR loci were analyzed by PCR followed PAGE and silver staining. RESULTS: The allele frequencies of the five X-STRs in Yunnan Nu population are in accordance with Hardy-Weinberg equilibrium. CONCLUSION: Five X-STRs loci of Nu population could be used in forensic identification.

Alleles↗

Regulatory context is a crucial part of gene function.

Information about the time and place of gene transcription, which until recently was only possible by extensive experimental analysis, can now be predicted through in silico analysis. Using the human RANTES/CCL5 promoter, we show that organizational features of promoters derived from promoter sequences contain information about the spatial and temporal 'functional context' of expression.

Animals↗

Paircomp, FamilyRelationsII and Cartwheel: tools for interspecific sequence comparison.

BACKGROUND: Comparative sequence analysis is an effective and increasingly common way to identify cis-regulatory regions in animal genomes. RESULTS: We describe three tools for comparative analysis of pairs of BAC-sized genomic regions. Paircomp is a tool that does windowed (ungapped) comparisons of two sequences and reports all matches above a set threshold. FamilyRelationsII is a graphical viewer for comparisons that enables interactive exploration of several different kinds of comparisons. Cartwheel is a Web site and compute-cluster management system used to execute and store comparisons for display by FamilyRelationsII. These tools are specialized for the discovery of cis-regulatory regions in animal genomes. All tools and their source code are freely available at http://family.caltech.edu/. CONCLUSION: These tools have been shown to effectively identify regulatory regions in echinoderms, mammals, and nematodes.

Algorithms↗

XenDB: full length cDNA prediction and cross species mapping in Xenopus laevis.

BACKGROUND: Research using the model system Xenopus laevis has provided critical insights into the mechanisms of early vertebrate development and cell biology. Large scale sequencing efforts have provided an increasingly important resource for researchers. To provide full advantage of the available sequence, we have analyzed 350,468 Xenopus laevis Expressed Sequence Tags (ESTs) both to identify full length protein encoding sequences and to develop a unique database system to support comparative approaches between X. laevis and other model systems. DESCRIPTION: Using a suffix array based clustering approach, we have identified 25,971 clusters and 40,877 singleton sequences. Generation of a consensus sequence for each cluster resulted in 31,353 tentative contig and 4,801 singleton sequences. Using both BLASTX and FASTY comparison to five model organisms and the NR protein database, more than 15,000 sequences are predicted to encode full length proteins and these have been matched to publicly available IMAGE clones when available. Each sequence has been compared to the KOG database and approximately 67% of the sequences have been assigned a putative functional category. Based on sequence homology to mouse and human, putative GO annotations have been determined. CONCLUSION: The results of the analysis have been stored in a publicly available database XenDB http://bibiserv.techfak.uni-bielefeld.de/xendb/. A unique capability of the database is the ability to batch upload cross species queries to identify potential Xenopus homologues and their associated full length clones. Examples are provided including mapping of microarray results and application of 'in silico' analysis. The ability to quickly translate the results of various species into 'Xenopus-centric' information should greatly enhance comparative embryological approaches.

Animals↗

Comparison of the small molecule metabolic enzymes of Escherichia coli and Saccharomyces cerevisiae.

The comparison of the small molecule metabolism pathways in Escherichia coli and Saccharomyces cerevisiae (yeast) shows that 271 enzymes are common to both organisms. These common enzymes involve 384 gene products in E. coli and 390 in yeast, which are between one half and two thirds of the gene products of small molecule metabolism in E. coli and yeast, respectively. The arrangement and family membership of the domains that form all or part of 374 E. coli sequences and 343 yeast sequences was determined. Of these, 70% consist entirely of homologous domains, and 20% have homologous domains linked to other domains that are unique to E. coli, yeast, or both. Over two thirds of the enzymes common to the two organisms have sequence identities between 30% and 50%. The remaining groups include 13 clear cases of nonorthologous displacement. Our calculations show that at most one half to two thirds of the gene products involved in small molecule metabolism are common to E. coli and yeast. We have shown that the common core of 271 enzymes has been largely conserved since the separation of prokaryotes and eukaryotes, including modifications for regulatory purposes, such as gene fusion and changes in the number of isozymes in one of the two organisms. Only one fifth of the common enzymes have nonhomologous domains between the two organisms. Around the common core very different extensions have been made to small molecule metabolism in the two organisms.

Databases, Genetic↗

Integrating alternative splicing detection into gene prediction.

BACKGROUND: Alternative splicing (AS) is now considered as a major actor in transcriptome/proteome diversity and it cannot be neglected in the annotation process of a new genome. Despite considerable progresses in term of accuracy in computational gene prediction, the ability to reliably predict AS variants when there is local experimental evidence of it remains an open challenge for gene finders. RESULTS: We have used a new integrative approach that allows to incorporate AS detection into ab initio gene prediction. This method relies on the analysis of genomically aligned transcript sequences (ESTs and/or cDNAs), and has been implemented in the dynamic programming algorithm of the graph-based gene finder EuGENE. Given a genomic sequence and a set of aligned transcripts, this new version identifies the set of transcripts carrying evidence of alternative splicing events, and provides, in addition to the classical optimal gene prediction, alternative optimal predictions (among those which are consistent with the AS events detected). This allows for multiple annotations of a single gene in a way such that each predicted variant is supported by a transcript evidence (but not necessarily with a full-length coverage). CONCLUSIONS: This automatic combination of experimental data analysis and ab initio gene finding offers an ideal integration of alternatively spliced gene prediction inside a single annotation pipeline.

Algorithms↗

[Genetic polymorphism of 6 short tandem repeat loci in Mongolian population of China].

OBJECTIVE: To clarify the distribution of genetic polymorphism of D3S1358, D13S317, D5S818, D6S1043, D2S1772, D7S3048 loci of the Mongolian population in Ximeng pastoral area and construct the relevant genetic database. METHODS: Multiplex PCR and polyacrylamide gel electrophoresis were used to investigate the polymorphism of 6 short tandem repeat (STR) loci in 286 individuals of the Mongolian population. RESULTS: In this study, 6, 9, 8, 11, 14, 11 alleles were observed at the 6 STR loci respectively. The genotypes distributions in Mongolian population were in accordance with Hardy-Weinberg equilibrium (P>0.05), the cumulative expected heterozygosities (H), discriminating probability (DP) and the polymorphism information contents (PIC) for the 6 loci were 0.9998, 09999, 0.9998 respectively. These data were compared with those of the Han population. The results showed there were significant difference in D3S1358, D13S317, D5S818, D2S1772, D7S3048 loci between the Mongolian population and Han population (P<0.05). However, no significant difference in D6S1043 locus was seen between the two populations (P>0.05). CONCLUSION: The results demonstrate that these 6 STR loci can serve as genetic marks and provide valuable data which are beneficial to studying the population genetics and ethnology.

China↗