Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

A hitchhiker's guide to expressed sequence tag (EST) analysis.

Expressed sequence tag (EST) sequencing projects are underway for numerous organisms, generating millions of short, single-pass nucleotide sequence reads, accumulating in EST databases. Extensive computational strategies have been developed to organize and analyse both small- and large-scale EST data for gene discovery, transcript and single nucleotide polymorphism analysis as well as functional annotation of putative gene products. We provide an overview of the significance of ESTs in the genomic era, their properties and the applications of ESTs. Methods adopted for each step of EST analysis by various research groups have been compared. Challenges that lie ahead in organizing and analysing the ever increasing EST data have also been identified. The most appropriate software tools for EST pre-processing, clustering and assembly, database matching and functional annotation have been compiled (available online from http://biolinfo.org/EST). We propose a road map for EST analysis to accelerate the effective analyses of EST data sets. An investigation of EST analysis platforms reveals that they all terminate prior to downstream functional annotation including gene ontologies, motif/pattern analysis and pathway mapping.

Animals↗

Generation of a database containing discordant intron positions in eukaryotic genes (MIDB).

MOTIVATION: Intron sliding is the relocation of intron-exon boundaries over short distances and is often also referred to as intron slippage or intron migration or intron drift. We have generated a database containing discordant intron positions in homologous genes (MIDB--Mismatched Intron DataBase). Discordant intron positions are those that are either closely located in homologous genes (within a window of 10 nucleotides) or an intron position that is present in one gene but not in any of its homologs. The MIDB database aims at systematically collecting information about mismatched introns in the genes from GenBank and organizing it into a form useful for understanding the genomics and dynamics of introns thereby helping understand the evolution of genes. RESULTS: Intron displacement or sliding is critically important for explaining the present distribution of introns among orthologous and paralogous genes. MIDB allows examining of intron movements and allows mapping of intron positions from homologous proteins onto a single sequence. The database is of potential use for molecular biologists in general and for researchers who are interested in gene evolution and eukaryotic gene structure. Partial analysis of this database allowed us to identify a few putative cases of intron sliding. AVAILABILITY: http://intron.bic.nus.edu.sg/midb/midb.html

Amino Acid Sequence↗

SIR: a simple indexing and retrieval system for biological flat file databases.

SUMMARY: SIR is a Simple Indexing and Retrieval tool for indexing and searching biological flat file databases. SIR is a cross-platform solution entirely written in Python. Since the package is very small and installation is trivial, this would be an ideal solution for database providers to provide a custom retrieval tool to access them. AVAILABILITY: The modules will be made available at http://www.EMBLHeidelberg.de/~chenna/PySAT/sir.html

Abstracting and Indexing↗

Condition specific transcription factor binding site characterization in Saccharomyces cerevisiae.

MOTIVATION: We demonstrate a computational process by which transcription factor binding sites can be elucidated using genome-wide expression and binding profiles. The profiles direct us to the intergenic locations likely to contain the promoter regions for a given factor. These sequences are multiply and locally aligned to give an anchor motif from which further characterization can take place. RESULTS: We present bases for and assumptions about the variability within these motifs which give rise to potentially more accurate motifs, capture complex binding sites built upon the basis motif, and eliminate the constraints of the currently employed promoter searching protocols. We also present a measure of motif quality based on the occurrence of the putative motifs in regions observed to contain the binding sites. The assumptions, motif generation, quality assessment and comparison allow the user as much control as their a priori knowledge allows. AVAILABILITY: IGRDB and the datasets mentioned herein are available at http://chipdb.wi.mit.edu/

Amino Acid Motifs↗

Gene structure identification with MyGV using cDNA evidence and protein homologs to improve ab initio predictions.

UNLABELLED: MyGV is an application to visualize (potentially genome-scale) gene structure annotation and prediction. The output of any external gene prediction program can be easily converted to a generalized format for input into MyGV. The application displays all input simultaneously in graphical representation, with a toggle option for a text-based view. Zooming capabilities allow detailed comparisons for specific genome locations. The tool is particularly helpful for refinement of ab initio predicted gene structures by spliced alignment with cDNA or protein homologs. AVAILABILITY: The program was written in Java and is freely available to non-commercial users by electronic download from http://bioinformatics.iastate.edu/bioinformatics2go/MyGV.

Animals↗

Predicting HIV drug resistance with neural networks.

MOTIVATION: Drug resistance is a very important factor influencing the failure of current HIV therapies. The ability to predict the drug resistance of HIV protease mutants may be useful in developing more effective and longer lasting treatment regimens. METHODS: The HIV resistance is predicted to two current protease inhibitors, Indinavir and Saquinavir. The problem was approached from two perspectives. First, a predictor was constructed based on the structural features of the HIV protease-drug inhibitor complex. A particular structure was represented by its list of contacts between the inhibitor and the protease. Next, a classifier was constructed based on the sequence data of various drug resistant mutants. In both cases, self-organizing maps were first used to extract the important features and cluster the patterns in an unsupervised manner. This was followed by subsequent labelling based on the known patterns in the training set. RESULTS: The prediction performance of the classifiers was measured by cross-validation. The classifier using the structure information correctly classified previously unseen mutants with an accuracy of between 60 and 70%. Several architectures were tested on the more abundant sequence data. The best single classifier provided an accuracy of 68% and a coverage of 69%. Multiple networks were then combined into various majority voting schemes. The best combination yielded an average of 85% coverage and 78% accuracy on previously unseen data. This is more than two times better than the 33% accuracy expected from a random classifier.

Algorithms↗

Evaluation of annotation strategies using an entire genome sequence.

MOTIVATION: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. RESULTS: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences.

Amino Acid Sequence↗

MeKE: discovering the functions of gene products from biomedical literature via sentence alignment.

MOTIVATION: Research on roles of gene products in cells is accumulating and changing rapidly, but most of the results are still reported in text form and are not directly accessible by computers. To expedite the progress of functional bioinformatics, it is, therefore, important to efficiently process large amounts of biomedical literature and transform the knowledge extracted into a structured format usable by biologists and medical researchers. Our aim was to develop an intelligent text-mining system that will extract from biomedical documents knowledge about the functions of gene products and thus facilitate computing with function. RESULTS: We have developed an ontology-based text-mining system to efficiently extract from biomedical literature knowledge about the functions of gene products. We also propose methods of sentence alignment and sentence classification to discover the functions of gene products discussed in digital texts. AVAILABILITY: http://ismp.csie.ncku.edu.tw/~yuhc/meke/

Biomedical Research↗

Data integration and visualization system for enabling conceptual biology.

MOTIVATION: Integration of heterogeneous data in life sciences is a growing and recognized challenge. The problem is not only to enable the study of such data within the context of a biological question but also more fundamentally, how to represent the available knowledge and make it accessible for mining. RESULTS: Our integration approach is based on the premise that relationships between biological entities can be represented as a complex network. The context dependency is achieved by a judicious use of distance measures on these networks. The biological entities and the distances between them are mapped for the purpose of visualization into the lower dimensional space using the Sammon's mapping. The system implementation is based on a multi-tier architecture using a native XML database and a software tool for querying and visualizing complex biological networks. The functionality of our system is demonstrated with two examples: (1) A multiple pathway retrieval, in which, given a pathway name, the system finds all the relationships related to the query by checking available metabolic pathway, transcriptional, signaling, protein-protein interaction and ontology annotation resources and (2) A protein neighborhood search, in which given a protein name, the system finds all its connected entities within a specified depth. These two examples show that our system is able to conceptually traverse different databases to produce testable hypotheses and lead towards answers to complex biological questions.

Computational Biology↗

A procedure for assessing GO annotation consistency.

MOTIVATION: The Gene Ontology (GO) is widely used to annotate molecular attributes of genes and gene products. Multiple groups undertaking functional annotations of genomes contribute their annotation sets to the GO database resource and these data are subsequently used in comparative functional analysis research. Although GO curators adhere to the same protocols and standards while assigning GO annotations, the specific procedure followed by each annotation group can vary. Since differences in application of annotation standards would dilute the effectiveness of comparative analysis, methods for assessing annotation consistency are essential. The development of methodologies that are broadly applicable for the assessment of GO annotation consistency is an important issue for the comparative genomics community. RESULTS: We have developed a methodology for assessing the consistency of GO annotations provided by different annotation groups. The method is completely general and can be applied to compare any two sets of GO annotations. This is the first attempt to assess cross-species GO annotation consistency. Our method compares annotation sets utilizing the hierarchical structure of the GO to compare GO annotations between orthologous gene pairs. The method produces a report on the annotation consistency and inconsistency for each orthologous pair. We present results obtained by comparing GO annotations for mouse and human gene sets. AVAILABILITY: The complete current MGI_GOA GO annotation consistency report is available online at http://www.spatial.maine.edu/~mdolan/

Animals↗

Robust classification modeling on microarray data using misclassification penalized posterior.

MOTIVATION: Genome-wide microarray data are often used in challenging classification problems of clinically relevant subtypes of human diseases. However, the identification of a parsimonious robust prediction model that performs consistently well on future independent data has not been successful due to the biased model selection from an extremely large number of candidate models during the classification model search and construction. Furthermore, common criteria of prediction model performance, such as classification error rates, do not provide a sensitive measure for evaluating performance of such astronomic competing models. Also, even though several different classification approaches have been utilized to tackle such classification problems, no direct comparison on these methods have been made. RESULTS: We introduce a novel measure for assessing the performance of a prediction model, the misclassification-penalized posterior (MiPP), the sum of the posterior classification probabilities penalized by the number of incorrectly classified samples. Using MiPP, we implement a forward step-wise cross-validated procedure to find our optimal prediction models with different numbers of features on a training set. Our final robust classification model and its dimension are determined based on a completely independent test dataset. This MiPP-based classification modeling approach enables us to identify the most parsimonious robust prediction models only with two or three features on well-known microarray datasets. These models show superior performance to other models in the literature that often have more than 40-100 features in their model construction. AVAILABILITY: Our MiPP software program is available at the Bioconductor website (http://www.bioconductor.org).

Algorithms↗

Mapping SNPs to protein sequence and structure data.

MOTIVATION: Data on both single nucleotide polymorphisms and disease-related mutations are being collected at ever-increasing rates. To understand the structural effects of missense mutations, we consider both classes under the term single amino acid polymorphisms (SAAPs) and we wish to map these to protein structure where their effects can be analyzed. Our initial aim therefore is to create a completely automatically maintained database of SAAPs mapped to individual residues in the Protein Data Bank (PDB) updated as new mutations or structures become available. RESULTS: We present an integrated pipeline for the automated mapping of SAAP data from HGVbase to individual PDB residues. Achieving this in a completely automated and reliable manner is a complex task. Data extracted from HGVbase are mapped to EMBL entries to confirm whether the mutation occurs in an exon and, if so, where in the sequence it occurs. From there we map to Swiss-Prot entries and thence to the PDB. AVAILABILITY: The resulting database may be accessed over the web at http://www.bioinf.org.uk/saap/ or http://acrmwww.biochem.ucl.ac.uk/saap/ CONTACT: a.martin@biochem.ucl.ac.uk.

Algorithms↗

The predictive power of the CluSTr database.

SUMMARY: The CluSTr database employs a fully automatic single-linkage hierarchical clustering method based on a similarity matrix. In order to compute the matrix, first all-against-all pair-wise comparisons between protein sequences are computed using the Smith-Waterman algorithm. The statistical significance of the similarity scores is then assessed using a Monte Carlo analysis, yielding Z-values, which are used to populate the matrix. This paper describes automated annotation experiments that quantify the predictive power and hence the biological relevance of the CluSTr data. The experiments utilized the UniProt data-mining framework to derive annotation predictions using combinations of InterPro and CluSTr. We show that this combination of data sources greatly increases the precision of predictions made by the data-mining framework, compared with the use of InterPro data alone. We conclude that the CluSTr approach to clustering proteins makes a valuable contribution to traditional protein classifications. AVAILABILITY: http://www.ebi.ac.uk/clustr/.

Algorithms↗

GeneCruiser: a web service for the annotation of microarray data.

SUMMARY: GeneCruiser is a web service allowing users to annotate their genomic data by mapping microarray feature identifiers to gene identifiers from databases, such as UniGene, while providing links to web resources, such as the UCSC Genome Browser. It relies on a regularly updated database that retrieves and indexes the mappings between microarray probes and genomic databases. Genes are identified using the Life Sciences Identifier standard. AVAILABILITY: GeneCruiser is freely available in the following forms: Web service and Web application, http://www.genecruiser.org; GenePattern, GeneCruiser access has been integrated into our microarray analysis platform, GenePattern. http://www.genepattern.org.

Animals↗

BlastXtract--a new way of exploring translated searches.

SUMMARY: Searches of translated, unannotated genomic DNA sequences against protein databases is a useful early-stage method for discovering protein homologues encoded by the sequence, but generates huge amounts of output data that quickly become impregnable. BlastXtract is a web-based tool for managing and visualizing results from large translated BLAST and FastA searches. It combines the speed and storage benefits of relational database management systems with an easy-to-use graphical navigation map, and greatly facilitates the early exploration of genomic sequence. AVAILABILITY: BlastXtract can be downloaded from http://bioinfo.ucc.ie/blastxtract/.

Amino Acid Sequence↗

Friend, an integrated analytical front-end application for bioinformatics.

UNLABELLED: Friend is a bioinformatics application designed for simultaneous analysis and visualization of multiple structures and sequences of proteins and/or DNA/RNA. The application provides basic functionalities, such as structure visualization, with different rendering and coloring, sequence alignment and simple phylogeny analysis, along with a number of extended features to perform more complex analyses of sequence structure relationships, including structural alignment of proteins, investigation of specific interaction motifs, studies of protein-protein and protein-DNA interactions and protein super-families. It is also useful for functional annotation of proteins, protein modeling and protein folding studies. Friend provides three levels of usage: (1) an extensive GUI for a scientist with no programming experience, (2) a command line interface for scripting for a scientist with some programming experience and (3) the ability to extend Friend with user written libraries for an experienced programmer. The application is linked and communicates with local and remote sequence and structure databases. AVAILABILITY: http://mozart.bio.neu.edu/friend.

Computational Biology↗

GenColors: accelerated comparative analysis and annotation of prokaryotic genomes at various stages of completeness.

SUMMARY: GenColors is a new web-based software/database system aimed at an improved and accelerated annotation of prokaryotic genomes, considering information on related genomes and making extensive use of genome comparison. It offers a seamless integration of data from ongoing sequencing projects and annotated genomic sequences obtained from GenBank. The genome comparison tools determine, for example, best-bidirectional hits, gene conservation, syntenies and gene core sets. Swiss-Prot/TrEMBL hits allow annotations in an effective manner. To further support the annotation base-specific quality data can also be displayed if available. With GenColors dedicated genome browsers containing a group of related genomes can be easily set up and maintained. It has been efficiently used for Borrelia garinii and is currently applied to various ongoing genome projects. AVAILABILITY: Detailed information on GenColors is available at http://gencolors.imb-jena.de. Online usage of GenColors-based genome browsers is the preferred application mode. The system is also available upon request for local installation.

Borrelia↗

MACiE: a database of enzyme reaction mechanisms.

SUMMARY: MACiE (mechanism, annotation and classification in enzymes) is a publicly available web-based database, held in CMLReact (an XML application), that aims to help our understanding of the evolution of enzyme catalytic mechanisms and also to create a classification system which reflects the actual chemical mechanism (catalytic steps) of an enzyme reaction, not only the overall reaction. AVAILABILITY: http://www-mitchell.ch.cam.ac.uk/macie/.

Catalysis↗