Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Evidence that plant-like genes in Chlamydia species reflect an ancestral relationship between Chlamydiaceae, cyanobacteria, and the chloroplast.

An unusually high proportion of proteins encoded in Chlamydia genomes are most similar to plant proteins, leading to proposals that a Chlamydia ancestor obtained genes from a plant or plant-like host organism by horizontal gene transfer. However, during an analysis of bacterial-eukaryotic protein similarities, we found that the vast majority of plant-like sequences in Chlamydia are most similar to plant proteins that are targeted to the chloroplast, an organelle derived from a cyanobacterium. We present further evidence suggesting that plant-like genes in Chlamydia, and other Chlamydiaceae, are likely a reflection of an unappreciated evolutionary relationship between the Chlamydiaceae and the cyanobacteria-chloroplast lineage. Further analyses of bacterial and eukaryotic genomes indicates the importance of evaluating organellar ancestry of eukaryotic proteins when identifying bacteria-eukaryote homologs or horizontal gene transfer and supports the proposal that Chlamydiaceae, which are obligate intracellular bacterial pathogens of animals, are not likely exchanging DNA with their hosts.

Animals↗

The MicrobesOnline Web site for comparative genomics.

At present, hundreds of microbial genomes have been sequenced, and hundreds more are currently in the pipeline. The Virtual Institute for Microbial Stress and Survival has developed a publicly available suite of Web-based comparative genomic tools (http://www.microbesonline.org) designed to facilitate multispecies comparison among prokaryotes. Highlights of the MicrobesOnline Web site include operon and regulon predictions, a multispecies genome browser, a multispecies Gene Ontology browser, a comparative KEGG metabolic pathway viewer, a Bioinformatics Workbench for in-depth sequence analysis, and Gene Carts that allow users to save genes of interest for further study while they browse. In addition, we provide an interface for genome annotation, which like all of the tools reported here, is freely available to the scientific community.

Animals↗

Large-scale protein annotation through gene ontology.

Recent progress in genomic sequencing, computational biology, and ontology development has presented an opportunity to investigate biological systems from a unique perspective, that is, examining genomes and transcriptomes through the multiple and hierarchical structure of Gene Ontology (GO). We report here our development of GO Engine, a computational platform for GO annotation, and analysis of the resultant GO annotations of human proteins. Protein annotation was centered on sequence homology with GO-annotated proteins and protein domain analysis. Text information analysis and a multiparameter cellular localization predictive tool were also used to increase the annotation accuracy, and to predict novel annotations. The majority of proteins corresponding to full-length mRNA in GenBank, and the majority of proteins in the NR database (nonredundant database of proteins) were annotated with one or more GO nodes in each of the three GO categories. The annotations of GenBank and SWISS-PROT proteins are available to the public at the GO Consortium web site.

Animals↗

Comparative gene prediction in human and mouse.

The completion of the sequencing of the mouse genome promises to help predict human genes with greater accuracy. While current ab initio gene prediction programs are remarkably sensitive (i.e., they predict at least a fragment of most genes), their specificity is often low, predicting a large number of false-positive genes in the human genome. Sequence conservation at the protein level with the mouse genome can help eliminate some of those false positives. Here we describe SGP2, a gene prediction program that combines ab initio gene prediction with TBLASTX searches between two genome sequences to provide both sensitive and specific gene predictions. The accuracy of SGP2 when used to predict genes by comparing the human and mouse genomes is assessed on a number of data sets, including single-gene data sets, the highly curated human chromosome 22 predictions, and entire genome predictions from ENSEMBL. Results indicate that SGP2 outperforms purely ab initio gene prediction methods. Results also indicate that SGP2 works about as well with 3x shotgun data as it does with fully assembled genomes. SGP2 provides a high enough specificity that its predictions can be experimentally verified at a reasonable cost. SGP2 was used to generate a complete set of gene predictions on both the human and mouse by comparing the genomes of these two species. Our results suggest that another few thousand human and mouse genes currently not in ENSEMBL are worth verifying experimentally.

Animals↗

The mammalian protein-protein interaction database and its viewing system that is linked to the main FANTOM2 viewer.

Here, we describe the development of a mammalian protein-protein interaction (PPI) database and of a PPI Viewer application to display protein interaction networks (http://fantom21.gsc.riken.go.jp/PPI/). In the database, we stored the mammalian PPIs identified through our PPI assays (internal PPIs), as well as those we extracted and processed (external PPIs) from publicly available data sources, the DIP and BIND databases and MEDLINE abstracts by using FACTS, a new functional inference and curation system. We integrated the internal and external PPIs into the PPI database, which is linked to the main FANTOM2 viewer. In addition, we incorporated into the PPI Viewer information regarding the luciferase reporter activity of internal PPIs and the data confidence of external PPIs; these data enable visualization and evaluation of the reliability of each interaction. Using the described system, we successfully identified several interactions of biological significance. Therefore, the PPI Viewer is a useful tool for exploring FANTOM2 clone-related protein interactions and their potential effects on signaling and cellular communication.

Animals↗

AraCyc: a biochemical pathway database for Arabidopsis.

AraCyc is a database containing biochemical pathways of Arabidopsis, developed at The Arabidopsis Information Resource (http://www.arabidopsis.org). The aim of AraCyc is to represent Arabidopsis metabolism as completely as possible with a user-friendly Web-based interface. It presently features more than 170 pathways that include information on compounds, intermediates, cofactors, reactions, genes, proteins, and protein subcellular locations. The database uses Pathway Tools software, which allows the users to visualize a bird's eye view of all pathways in the database down to the individual chemical structures of the compounds. The database was built using Pathway Tools' Pathologic module with MetaCyc, a collection of pathways from more than 150 species, as a reference database. This initial build was manually refined and annotated. More than 20 plant-specific pathways, including carotenoid, brassinosteroid, and gibberellin biosyntheses have been added from the literature. A list of more than 40 plant pathways will be added in the coming months. The quality of the initial, automatic build of the database was compared with the manually improved version, and with EcoCyc, an Escherichia coli database using the same software system that has been manually annotated for many years. In addition, a Perl interface, PerlCyc, was developed that allows programmers to access Pathway Tools databases from the popular Perl language. AraCyc is available at the tools section of The Arabidopsis Information Resource Web site (http://www.arabidopsis.org/tools/aracyc).

Arabidopsis↗

Munich information center for protein sequences plant genome resources: a framework for integrative and comparative analyses 1(W).

With several plant genomes sequenced, the power of comparative genome analysis can now be applied. However, genome-scale cross-species analyses are limited by the effort for data integration. To develop an integrated cross-species plant genome resource, we maintain comprehensive databases for model plant genomes, including Arabidopsis (Arabidopsis thaliana), maize (Zea mays), Medicago truncatula, and rice (Oryza sativa). Integration of data and resources is emphasized, both in house as well as with external partners and databases. Manual curation and state-of-the-art bioinformatic analysis are combined to achieve quality data. Easy access to the data is provided through Web interfaces and visualization tools, bulk downloads, and Web services for application-level access. This allows a consistent view of the model plant genomes for comparative and evolutionary studies, the transfer of knowledge between species, and the integration with functional genomics data.

Computational Biology↗

MetaCyc and AraCyc. Metabolic pathway databases for plant research.

MetaCyc (http://metacyc.org) contains experimentally determined biochemical pathways to be used as a reference database for metabolism. In conjunction with the Pathway Tools software, MetaCyc can be used to computationally predict the metabolic pathway complement of an annotated genome. To increase the breadth of pathways and enzymes, more than 60 plant-specific pathways have been added or updated in MetaCyc recently. In contrast to MetaCyc, which contains metabolic data for a wide range of organisms, AraCyc is a species-specific database containing only enzymes and pathways found in the model plant Arabidopsis (Arabidopsis thaliana). AraCyc (http://arabidopsis.org/tools/aracyc/) was the first computationally predicted plant metabolism database derived from MetaCyc. Since its initial computational build, AraCyc has been under continued curation to enhance data quality and to increase breadth of pathway coverage. Twenty-eight pathways have been manually curated from the literature recently. Pathway predictions in AraCyc have also been recently updated with the latest functional annotations of Arabidopsis genes that use controlled vocabulary and literature evidence. AraCyc currently features 1,418 unique genes mapped onto 204 pathways with 1,156 literature citations. The Omics Viewer, a user data visualization and analysis tool, allows a list of genes, enzymes, or metabolites with experimental values to be painted on a diagram of the full pathway map of AraCyc. Other recent enhancements to both MetaCyc and AraCyc include implementation of an evidence ontology, which has been used to provide information on data quality, expansion of the secondary metabolism node of the pathway ontology to accommodate curation of secondary metabolic pathways, and enhancement of the cellular component ontology for storing and displaying enzyme and pathway locations within subcellular compartments.

4-Hydroxyphenylpyruvate Dioxygenase↗

Effective electron-density map improvement and structure validation on a Linux multi-CPU web cluster: The TB Structural Genomics Consortium Bias Removal Web Service.

Anticipating a continuing increase in the number of structures solved by molecular replacement in high-throughput crystallography and drug-discovery programs, a user-friendly web service for automated molecular replacement, map improvement, bias removal and real-space correlation structure validation has been implemented. The service is based on an efficient bias-removal protocol, Shake&wARP, and implemented using EPMR and the CCP4 suite of programs, combined with various shell scripts and Fortran90 routines. The service returns improved maps, converted data files and real-space correlation and B-factor plots. User data are uploaded through a web interface and the CPU-intensive iteration cycles are executed on a low-cost Linux multi-CPU cluster using the Condor job-queuing package. Examples of map improvement at various resolutions are provided and include model completion and reconstruction of absent parts, sequence correction, and ligand validation in drug-target structures.

Apolipoproteins E↗

Functional genomics databases on the web.

Experiments involving high-throughput methods for measuring transcripts, proteins and metabolites constitute the area of functional genomics. These experiments are highly context dependent and require much more detail about the experimental design, sample and protocols used than in genomics. Functional genomics databases are needed that follow established and emerging standards. Functional genomic databases are not yet very common; however, there are a few focused on microbial genomes and a couple integrative systems are available for setting up functional genomics databases.

Computational Biology↗

A combined experimental and computational strategy to define protein interaction networks for peptide recognition modules.

Peptide recognition modules mediate many protein-protein interactions critical for the assembly of macromolecular complexes. Complete genome sequences have revealed thousands of these domains, requiring improved methods for identifying their physiologically relevant binding partners. We have developed a strategy combining computational prediction of interactions from phage-display ligand consensus sequences with large-scale two-hybrid physical interaction tests. Application to yeast SH3 domains generated a phage-display network containing 394 interactions among 206 proteins and a two-hybrid network containing 233 interactions among 145 proteins. Graph theoretic analysis identified 59 highly likely interactions common to both networks. Las17 (Bee1), a member of the Wiskott-Aldrich Syndrome protein (WASP) family of actin-assembly proteins, showed multiple SH3 interactions, many of which were confirmed in vivo by coimmunoprecipitation.

Algorithms↗

Genome sequence-based fluorescent amplified fragment length polymorphism of Campylobacter jejuni, its relationship to serotyping, and its implications for epidemiological analysis.

The published genome sequence of Campylobacter jejuni strain NCTC 11168 was used to model an accurate and highly reproducible fluorescent amplified fragment length polymorphism (FAFLP) analysis. Predicted and experimentally observed amplified fragments (AFs) generated with the primer pair HindIII+A and HhaI+A were compared. All but one of the 61 predicted AFs were reproducibly detected, and no unpredicted fragments were amplified. This FAFLP analysis was used to genotype 74 C. jejuni strains belonging to the nine heat-stable (HS) serotypes most prevalent in human disease in England and Wales. The 74 C. jejuni strains exhibited 60 FAFLP profiles, and cluster analysis of them yielded a radial tree showing genetic relationships between and within 13 major clusters. Some clusters were related, and others were unrelated, to a single HS serotype. For example, all strains belonging to serotypes HS6 and HS19 grouped into corresponding single genotypic clusters, while strains of serotypes HS11 and HS18 each grouped into two genotypic clusters. Strains of HS50, the most prevalent serotype infecting humans, were found both in one large (multiserotype) cluster complex and dispersed throughout the tree. The strain genotypes within each FAFLP cluster were characterized by a particular combination of AFs, and among the cluster there were additional differential AFs. Identification of such AFs could act as a search tool to look for potential associations with disease or animal hosts, when applied to large number of human isolates. Genome-sequence based FAFLP, thus, has the potential to establish a genetic database for epidemiological investigations of Campylobacter.

Animals↗

BIBI, a bioinformatics bacterial identification tool.

BIBI was designed to automate DNA sequence analysis for bacterial identification in the clinical field. BIBI relies on the use of BLAST and CLUSTAL W programs applied to different subsets of sequences extracted from GenBank. These sequences are filtered and stored in a new database, which is adapted to bacterial identification.

Bacteria↗

A nonessential African swine fever virus gene UK is a significant virulence determinant in domestic swine.

Sequence analysis of the right variable genomic region of the pathogenic African swine fever virus (ASFV) isolate E70 revealed a novel gene, UK, that is immediately upstream from the previously described ASFV virulence-associated gene NL-S (L. Zsak, Z. Lu, G. F. Kutish, J. G. Neilan, and D. L. Rock, J. Virol. 70:8865-8871, 1996). UK, transcriptionally oriented toward the right end of the genome, predicts a protein of 96 amino acids with a molecular mass of 10.7 kDa. Searches of genetic databases did not find significant similarity between UK and other known genes. Sequence analysis of the UK genes from several pathogenic ASFVs from Europe, the Caribbean, and Africa demonstrated that this gene was highly conserved among diverse pathogenic isolates, including those from both tick and pig sources. Polyclonal antibodies raised against the UK protein specifically precipitated a 15-kDa protein from ASFV-infected macrophage cell cultures as early as 2 h postinfection. A recombinant UK gene deletion mutant, deltaUK, and its revertant, UK-R, were constructed from the E70 isolate to study gene function. Although deletion of UK did not affect the growth characteristics of the virus in macrophage cell cultures, deltaUK exhibited reduced virulence in infected pigs. While mortality among parental E70- or UK-R-infected animals was 100%, all deltaUK-infected pigs survived infection. Fever responses were comparable in E70-, UK-R-, and deltaUK-infected groups; however, deltaUK-infected animals exhibited significant, 100- to 1,000-fold, reductions in viremia titers. These data indicate that the highly conserved UK gene of ASFV, while being nonessential for growth in macrophages in vitro, is an important viral virulence determinant for domestic pigs.

African Swine Fever Virus↗

Optic disc anomalies and frontonasal dysplasia.

AIMS: To document the optic disc abnormalities in patients with frontonasal dysplasia in association with basal encephalocele. METHODS: Names and hospital numbers of patients with midline clefts were obtained from the ophthalmology and genetics database. Six patients were identified who had the following common findings: midline facial cleft with midline cleft lip and palate; hypertelorism; absent corpus callosum; basal (sphenoethmoidal) encephalocele; and pituitary deficiency (five out of six cases). Ophthalmic examination was performed with fundal photography where possible. RESULTS: Two patients had unilateral and one a bilateral peripapillary staphyloma. Two patients had bilateral optic disc hypoplasia and one appeared to have a peripapillary staphyloma in one eye and a morning glory disc in the other. CONCLUSION: Optic disc abnormalities were found in all patients with this constellation of clinical findings. This association appears to represent a distinct subgroup within the spectrum of frontonasal dysplasia. The presence of midline facial anomalies and any dysplastic disc should alert the physician as to the presence of an encephalocele.

Abnormalities, Multiple↗

A comprehensive whole genome bacterial phylogeny using correlated peptide motifs defined in a high dimensional vector space.

As whole genome sequences continue to expand in number and complexity, effective methods for comparing and categorizing both genes and species represented within extremely large datasets are required. Methods introduced to date have generally utilized incomplete and likely insufficient subsets of the available data. We have developed an accurate and efficient method for producing robust gene and species phylogenies using very large whole genome protein datasets. This method relies on multidimensional protein vector definitions supplied by the singular value decomposition (SVD) of a large sparse data matrix in which each protein is uniquely represented as a vector of overlapping tetrapeptide frequencies. Quantitative pairwise estimates of species similarity were obtained by summing the protein vectors to form species vectors, then determining the cosines of the angles between species vectors. Evolutionary trees produced using this method confirmed many accepted prokaryotic relationships. However, several unconventional relationships were also noted. In addition, we demonstrate that many of the SVD-derived right basis vectors represent particular conserved protein families, while many of the corresponding left basis vectors describe conserved motifs within these families as sets of correlated peptides (copeps). This analysis represents the most detailed simultaneous comparison of prokaryotic genes and species available to date.

Amino Acid Motifs↗