Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “bioinformatic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

The EMBL Nucleotide Sequence Database: major new developments.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/) incorporates, organizes and distributes nucleotide sequences from all available public sources. The database is located and maintained at the European Bioinformatics Institute (EBI) near Cambridge, UK. In an international collaboration with DDBJ (Japan) and GenBank (USA), data are exchanged amongst the collaborating databases on a daily basis to achieve optimal synchronization. Webin is the preferred web-based submission system for individual submitters, while automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO). Database releases are produced quarterly. Network services allow free access to the most up-to-date data collection via FTP, Email and World Wide Web interfaces. EBI's Sequence Retrieval System (SRS) integrates and links the main nucleotide and protein databases plus many other specialized molecular biology databases. For sequence similarity searching, a variety of tools (e.g. Fasta, BLAST) are available which allow external users to compare their own sequences against the latest data in the EMBL Nucleotide Sequence Database and SWISS-PROT. All resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗

[Correction of five different types of errors of model REFSEQs appeared in NCBI human gene database only by using two novel human genes C17orf32 and ZNF362].

Found that there exist many mistakes in the REFSEQ issued in the genome annotation project of NCBI, the result of which indicates that people be cautious in using REFSEQ database in NCBI. By adopting the technical route combining bioinformatics analysis and experimental verification, through the comparison of the cloned genes in the non-redundant database, we found that there were many mistakes in the computer annotation human genome coding sequences that were issued on the internet. First we quoted nine wrong types of novel human genes anticipated by NCBI GENOME Annotation Project. Here we give one example in detail: (1) Comparison of the sequences between novel human gene C17orf32 and hypothetical human gene LOC124919. LOC123722 is a modified sequence of C17orf32 cDNA with an inserted G between 406 -407 nucleotides. The base G in the 401 position of LOC123722 cDNA is a redundant insert, which causes a reading frame shift in the translation of an alternative protein. This inserted G has not been found in our experimental clone, and is fully rejected by human EST alignment, and is shown as a redundance by genomic GT/AG organization analysis. (2) Comparison of the sequences between novel human gene C17orf32 and hypothetical human gene LOC147007. C17orf32 gene (ORF from 31 to 657 nucleotides) is located on human chromosome 17(Accession No. NT_010808.7), and is only linked with a hypothetical human gene LOC147007 (ORF from 55 to 435 nucleotides) at present. This hypothetical human gene sequence has not been verified by experiment, and is a wrong form of our verified C17orf32 gene. The full-length 1 679 bp cDNA sequence of C17orf32 exhibits overall homology to that of LOC147007 of 625 bp mRNA, with matching percentage of 37% in 36% of total window over the full-length nucleotide, especially 121 approximately 366 bp of LOC147007 is just the same as 316 approximately 561 bp of C17orf32. Thus, the 126 aa protein encoded by XP_097165 of LOC147007 exhibits overall homology to the 208 aa protein encoded by C17orf32, with matching percentage of 50% in 48% of total window over the full-length protein, especially 23 approximately 104 aa of XP_097165 is just the same as 96 approximately 177 aa of C17orf32 protein. Both flanking regions of LOC147007 outside the same ORF central part are wrong assembly of non-relative cDNA. In addition, we have in silico cloned a novel mouse gene, ORF32 (open reading frame 32) with TPA accession number of BK000258, which is the mouse ortholog of human C17orf32. Our strategy is helpful in both finding out more novel human genes and correcting the mistakes in the REFSEQs issued by NCBI genome annnotation project. For example, we adopted the gene anticipating method, through automatic calculation and analysis, anticipated two modes reference sequences (LOC124919 and LOC147007) from NCBI contig NT_ 010808. Both of them should be C17orf32, but the fact is that both of them are various wrong forms of C17orf32, respectively are the first type and second type of mistakes. Another example, we adopted gene anticipation method, through automatic calculation and analysis, anticipated three modes reference sequences (LOC14907, LOC200084 and LOC91126) from NCBI contig NT_004511 which really are one type of gene of ZNF362, but submitted three different wrong forms of ZNF362, respectively are: the fourth, fifth, and seventh type of mistakes. We can correct or avoid the currently wrong human genome coding sequence by using in silico clone and combining experimental verification. People should be cautious in treating the computer's annotation which may exist all type of wrong human genome coding sequences. The correct identification and annotation of the novel human genes still remain to be a long and arduous task.

Amino Acid Sequence↗

Proteome profiling of human epithelial ovarian cancer cell line TOV-112D.

A proteome profiling of the epithelial ovarian cancer cell line TOV-112D was initiated as a protein expression reference in the study of ovarian cancer. Two complementary proteomic approaches were used in order to maximise protein identification: two-dimensional gel electrophoresis (2DE) protein separation coupled to matrix assisted laser desorption/ionisation time-of-flight mass spectrometry (MALDI-TOF MS) and one-dimensional gel electrophoresis (1DE) coupled to liquid-chromatography tandem mass spectrometry (LC MS/MS). One hundred and seventy-two proteins have been identified among 288 spots selected on two-dimensional gels and a total of 579 proteins were identified with the 1DE LC MS/MS approach. This proteome profiling covers a wide range of protein expression and identifies several proteins known for their oncogenic properties. Bioinformatics tools were used to mine databases in order to determine whether the identified proteins have previously been implicated in pathways associated with carcinogenesis or cell proliferation. Indeed, several of the proteins have been reported to be specific ovarian cancer markers while others are common to many tumorigenic tissues or proliferating cells. The diversity of proteins found and their association with known oncogenic pathways validate this proteomic approach. The proteome 2D map of the TOV-112D cell line will provide a valuable resource in studies on differential protein expression of human ovarian carcinomas while the 1DE LC MS/MS approach gives a picture of the actual protein profile of the TOV-112D cell line. This work represents one of the most complete ovarian protein expression analysis reports to date and the first comparative study of gene expression profiling and proteomic patterns in ovarian cancer.

Cell Line, Transformed↗

Distinct mutational landscapes for germline and somatic cancer variants in forty tumor suppressor genes.

Germline and somatic cancer variants in tumor suppressor genes (TSGs) share loss-of-function mechanisms, but studies of a few genes (DICER1 and CEBPA) have demonstrated differences in variant consequence and location. To systematically assess whether TSGs display distinct mutational patterns, we leveraged large public genetic databases and compared 32,941 high-quality pathogenic/likely pathogenic (P/LP) germline variants in ClinVar, with 12,907 oncogenic/likely oncogenic (O/LO) somatic tumor variants from cBioPortal across 40 TSGs. Only 3,863 (9.2%) variants were shared. Eighteen TSGs showed significantly different distributions of variant occurrences by molecular consequence, replicated with non-overlapping somatic data from the COSMIC database (chi-squared tests, false discovery rate = 5%). DICER1, TP53, and SMAD4 displayed excess somatic missense events, while nine TSGs (e.g., RB1 and APC) contained excess somatic stop-gain events throughout the coding sequence. Analysis by tumor type revealed excess stop-gain events in tissues exposed to environmental mutagens with corresponding mutation signatures. For several TSGs (WT1), germline variants predispose to tumors (Wilms' tumor) distinct from the majority source of somatic data (myeloid leukemia). Germline and somatic events are also distributed unevenly across cDNA locations, with 103 regions of preferential clustering in 39 TSGs (78 somatic and 25 germline). Twenty somatic clusters contained recurring frameshifts in homopolymer runs, many in tumors with microsatellite instability. Germline clusters contain more germline-exclusive variants, some driving non-cancer phenotypes reflecting genetic pleiotropy. Altogether, germline and somatic variants of TSGs represent unique sets with substantially different patterns shaped by selection pressures from gene-specific and somatic mutational mechanisms. Characterizing these distinctions enables more accurate clinical interpretation of TSG variants.

Humans↗

Characterisation and expression of an Hsp70 gene from Parastrongyloides trichosuri.

Parastrongyloides trichosuri is a nematode parasite of Australian brushtail possums that has an alternative free-living life cycle which can be readily maintained indefinitely in a laboratory setting. The ability to maintain this parasite in a free-living cycle and induce it to parasitism at the free-living L1 stage makes this an excellent model for the study of genes associated with parasitism. A 70kD protein from infective larvae of P. trichosuri that appears to be immunogenic in infected possums has been identified as a heat shock protein (Hsp)70 homologue. The complete gene for Pt-Hsp70 was cloned and sequenced. The protein encoded by the Pt-Hsp70 gene is the likely orthologue of the Caenorhabditis elegans protein, Hsp70A, also known as hsp-1. Reverse transcriptase-PCR data indicate that Pt-Hsp70 (designated Pt-hsp-1) is expressed at readily detectable levels in all developmental stages of both the parasitic and free-living P. trichosuri life cycles and the promoter is mildly inducible by heat shock. Bioinformatic analysis of expressed sequence tag databases indicates that C. eleganshsp-1 homologues, together with C. eleganshsp-3 homologues, are the predominant members of the Hsp70 superfamily that are normally expressed in parasitic stages of the Strongyloididae family. Promoter fusions to a beta-galactosidase coding sequence were prepared and introduced into wild type C. elegans to produce transgenic nematodes. Reporter gene expression was clearly present within embryonic cells and within intestinal cells of larval and adult stages. Thus, the expression of the Pt-hsp-1 promoter within P. trichosuri and transgenic C. elegans appears similar to the known expression of C. elegans hsp-1. This promoter should be of value in efforts to develop genetic manipulation tools for P. trichosuri.

Amino Acid Sequence↗

Characterization of a small plasmid (pMBCP) from bovine Pseudomonas pickettii that confers cadmium resistance.

This is the first report of isolation of Pseudomonas pickettii from a normal adult bovine duodenum. This organism was one of several bacteria isolated as part of a study to examine cadmium resistance genes (cad(r)) for use in generating transgenic plants to reclaim cadmium-contaminated soils in Kansas. P. pickettii containing a plasmid of 2.2kb (designated pMBCP) grew in Luria-Bertani broth and agar containing up to 800 microM of cadmium chloride and was resistant to 16 antibiotics. Curing the organism of plasmid revealed that antibiotic resistances were not plasmid-mediated. Low-level cadmium resistance was conferred by the plasmid because uncured organism grew significantly better (P<0.05) at 55 microM compared to cured organism. Both plasmid and chromosomal DNA were probed by DNA-DNA hybridization for the presence of known cadmium resistance genes (cadA, cadC, and cadD from Gram-positive (Staphylococcus aureus), but none were detected. The plasmid had one restriction site each for BamHI, PstI, SmaI, and XhoI; two sites each for HincII, SacI, and SphI; and multiple sites for AluI and XcmI. DNA sequence analyses of the cloned and original plasmids showed a GC content of greater than 60% and no homology to any published sequences in the GenBank, European Bioinformatics Institute, or Japanese Genome Net databases. The DNA sequence is contained in GenBank accession number AF144733. Thus, pMBCP offers low-level cadmium resistance to P. picketttii.

Animals↗

Identification, characterization and expression analysis of a new fibrillar collagen gene, COL27A1.

The fibrillar collagens provide structural scaffolding and strength to the extracellular matrices of connective tissues. We identified a partial sequence of a new fibrillar collagen gene in the NCBI databases and completed the sequence with bioinformatic approaches and 5' RACE. This gene, designated COL27A1, is approximately 156 kbp long and has 61 exons located on chromosome 9q32-33. The homologous mouse gene is located on chromosome 4. The gene encodes amino- and carboxyl-terminal propeptides similar to those in the 'minor' fibrillar collagens. The triple-helical domain is, however, shorter and contains 994 amino acids with two imperfections of the Gly-Xaa-Yaa repeat pattern. There were three sites of alternative RNA splicing, only one of which led to the intact mRNA that encodes this full-length collagen proalpha chain. Phylogenetic analyses indicated that COL27A1 forms a clade with COL24A1 that is distinct from the two known lineages of fibrillar collagens. Expression analyses of the mouse col27a1 gene demonstrated high expression in cartilage, the eye and ear, but also in lung and colon. It is likely that the major protein product of COL27A1, proalpha1(XXVII), is a component of the extracellular matrices of cartilage and these other tissues. Study of this collagen should yield insights into normal chondrogenesis, and provide clues to the pathogenesis of some chondrodysplasias and disorders of other tissues in which this gene is expressed.

Alternative Splicing↗

Bioinformatics and mass spectrometry for microorganism identification: proteome-wide post-translational modifications and database search algorithms for characterization of intact H. pylori.

MALDI-TOF mass spectrometry has been coupled with Internet-based proteome database search algorithms in an approach for direct microorganism identification. This approach is applied here to characterize intact H. pylori (strain 26695) Gram-negative bacteria, the most ubiquitous human pathogen. A procedure for including a specific and common posttranslational modification, N-terminal Met cleavage, in the search algorithm is described. Accounting for posttranslational modifications in putative protein biomarkers improves the identification reliability by at least an order of magnitude. The influence of other factors, such as number of detected biomarker peaks, proteome size, spectral calibration, and mass accuracy, on the microorganism identification success rate is illustrated as well.

Algorithms↗

The bestrophin family of anion channels: identification of prokaryotic homologues.

The human disease protein, Bestrophin-1, associated with vitelliform macular dystrophy, has recently been shown to be an integral membrane anion channel-forming protein. In this study we have recovered all bestrophin homologues from the NCBI database and analyzed their sequences using bioinformatic approaches. Eukaryotic homologues were found in animals and fungi but not in plants or protozoans, and prokaryotic homologues distantly related to the eukaryotic proteins, were identified in certain Gram-negative bacterial kingdoms but not in Gram-positive bacteria or archaea. Our analyses suggest a uniform 4 TMS topology for most of these homologues with regions of conservation overlapping and preceding the odd numbered TMSs and overlapping and following the even numbered TMSs. Well-conserved motifs were identified in both the eukaryotic and the prokaryotic homologues, and these proved to overlap, suggesting common structural and functional properties. Phylogenetic analyses revealed that the eukaryotic proteins cluster according to organismal type, and that the prokaryotic proteins sometimes (but not always) do so. This suggests that eukaryotic paralogues arose exclusively by recent gene duplication events although both early and late gene duplication events occurred in prokaryotes.

Animals↗

The Web as an educational tool for/in learning/teaching bioinformatics statistics.

Statistics provides essential tool in Bioinformatics to interpret the results of a database search or for the management of enormous amounts of information provided from genomics, proteomics and metabolomics. The goal of this project was the development of a software tool that would be as simple as possible to demonstrate the use of the Bioinformatics statistics. Computer Simulation Methods (CSMs) developed using Microsoft Excel were chosen for their broad range of applications, immediate and easy formula calculation, immediate testing and easy graphics representation, and of general use and acceptance by the scientific community. The result of these endeavours is a set of utilities which can be accessed from the following URL: http://gmein.uib.es/bioinformatica/statistics. When tested on students with previous coursework with traditional statistical teaching methods, the general opinion/overall consensus was that Web-based instruction had numerous advantages, but traditional methods with manual calculations were also needed for their theory and practice. Once having mastered the basic statistical formulas, Excel spreadsheets and graphics were shown to be very useful for trying many parameters in a rapid fashion without having to perform tedious calculations. CSMs will be of great importance for the formation of the students and professionals in the field of bioinformatics, and for upcoming applications of self-learning and continuous formation.

Computational Biology↗

Mitogen-activated protein kinase pathway was significantly activated in human bronchial epithelial cells by nicotine.

Nicotine is potentially associated with the onset of chronic obstructive pulmonary disease (COPD) and lung cancer. To gain insights into the molecular mechanism underlying such nicotine-induced conditions, microarray- bioinformatics analysis was carried out in the present study to explore the gene expression profiles in human bronchial epithelial cells (HBECs) treated with 5 microM nicotine for 4, 8, and 10 h. Of 1,800 assessed genes overall, 260 (14.4%) were upregulated and 17 (0.9%) down regulated significantly. Gene ontology analysis demonstrated that most of the differentially expressed genes belonged to the category of molecular function, especially to the subcategories of enzyme activity. The integration of obtained information with bioinformatics tools in DAVID and KEGG databases indicated that the greatest number of overexpressed genes was involved in mitogen-activated protein kinase (MAPK) pathway. Membrane array analysis subsequently suggested that both extracellular signal-regulated kinase (ERK) 1/2 and c-Jun-NH(2)-terminal kinase (JNK) signalings but not p38 MAPK signaling were activated in response to nicotine. Pretreatment of HBECs with specific inhibitors against ERK 1/2 and JNK but not p38 could significantly inhibit nicotine-induced interleukin- 8 production. These results suggest that MAPK pathway may mediate the effect of nicotine through ERK 1/2 and JNK but not p38 in HBECs treated with nicotine.

Base Sequence↗

Early bioinformatics: the birth of a discipline--a personal view.

MOTIVATION: The field of bioinformatics has experienced an explosive growth in the last decade, yet this 'new' field has a long history. Some historical perspectives have been previously provided by the founders of this field. Here, we take the opportunity to review the early stages and follow developments of this discipline from a personal perspective. RESULTS: We review the early days of algorithmic questions and answers in biology, the theoretical foundations of bioinformatics, the development of algorithms and database resources and finally provide a realistic picture of what the field looked like from a resources and finally provide a realistic picture of what the field looked like from a practitioner's viewpoint 10 years ago, with a perspective for future developments.

Algorithms↗

ER-associated protein degradation is a common mechanism underpinning numerous monogenic diseases including Robinow syndrome.

Correct folding of nascent polypeptide chains within the ER is critical for function, assembly into multi-subunit complexes and trafficking through the exocytic pathway for secretory and cell surface proteins. This process is rather inefficient, and a substantial proportion of nascent polypeptides is rejected by an ER quality control system and targeted for degradation. In some cases, only a minor fraction of nascent chains is correctly folded, and the smallest alteration to polypeptide primary structure (i.e. point mutation) can result in the complete loss of function with inherent pathological consequences; cystic fibrosis and emphysema result from such mutations. We have taken a bioinformatic approach to parse a large database of known disease susceptibility genes for candidates whose disease-associated alleles are likely prone to misfolding in the ER. Surprisingly, we find that proteins with ER-targeting signals are over represented in this database when compared with all predicted proteins in the human genome (45 versus 30%). We selected a subgroup of proteins that were positive for both an ER-targeting signal and a membrane-anchoring domain and thereby identified several ER-associated degradation diseases candidates. To determine whether our analysis had identified new ER-degradation substrates, we established that ER retention is indeed the mechanism underlying Robinow syndrome (RRS), one of the identified candidates. Specifically, mutant alleles of ROR2 that are associated with RRS are retained within the ER, whereas wild-type and non-pathogenic alleles are exported to the plasma membrane. These data both uncover a major pathogenic factor for RRS and indicate that misfolding of secretory proteins is likely to significantly contribute to human disease and morbidity.

Abnormalities, Multiple↗

The EMBL nucleotide sequence database.

The European Molecular Biology Laboratory (EMBL) Nucleotide Sequence Database (http://www.ebi.ac. uk/embl/index.html ) is maintained at the European Bioinformatics Institute (EBI) in an international collaboration with the DNA Data Bank of Japan (DDBJ) and GenBank (USA). Data is exchanged amongst the collaborative databases on a daily basis. The major contributors to the EMBL database are individual authors and genome project groups. WEBIN is the preferred web-based submission system for individual submitters, whilst automatic procedures allow incorporation of sequence data from large-scale genome sequencing centres and from the European Patent Office (EPO). Database releases are produced quarterly. Network services allow free access to the most up-to-date data collection via Internet and WWW interfaces. EBI's Sequence Retrieval System (SRS) is a network browser for databanks in molecular biology, integrating and linking the main nucleotide and protein databases plus many specialised databases. For sequence similarity searching a variety of tools (e.g., BLITZ, FASTA, BLAST) are available which allow external users to compare their own sequences against the most currently available data in the EMBL Nucleotide Sequence Database and SWISS-PROT.

Classification↗

The European Bioinformatics Institute's data resources.

As the amount of biological data grows, so does the need for biologists to store and access this information in central repositories in a free and unambiguous manner. The European Bioinformatics Institute (EBI) hosts six core databases, which store information on DNA sequences (EMBL-Bank), protein sequences (SWISS-PROT and TrEMBL), protein structure (MSD), whole genomes (Ensembl) and gene expression (ArrayExpress). But just as a cell would be useless if it couldn't transcribe DNA or translate RNA, our resources would be compromised if each existed in isolation. We have therefore developed a range of tools that not only facilitate the deposition and retrieval of biological information, but also allow users to carry out searches that reflect the interconnectedness of biological information. The EBI's databases and tools are all available on our website at www.ebi.ac.uk.

Animals↗

RNA/DNA Binding Protein TDP43 Regulates DNA Mismatch Repair Genes with Implications for Genome Stability.

TDP43 is an RNA/DNA binding protein increasingly recognized for its role in neurodegenerative conditions, including amyotrophic lateral sclerosis and frontotemporal dementia (FTD). As characterized by its aberrant nuclear export and cytoplasmic aggregation, TDP43 proteinopathy is a hallmark feature in over 95% of ALS/FTD cases, leading to the formation of detrimental cytosolic aggregates and a reduction in nuclear functionality within neurons. Building on our prior work linking TDP43 proteinopathy to the accumulation of DNA double-strand breaks (DSBs) in neurons, the present investigation uncovers a novel regulatory relationship between TDP43 and DNA mismatch repair (MMR) gene expressions. Here, we show that TDP43 depletion or overexpression directly affects the expression of key MMR genes. Alterations include MLH1, MSH2, MSH3, MSH6, and PMS2 levels across various primary cell lines, independent of their proliferative status. Our results specifically establish that TDP43 selectively influences the expression of MLH1 and MSH6 by influencing their alternative transcript splicing patterns and stability. We furthermore find aberrant MMR gene expression is linked to TDP43 proteinopathy in two distinct ALS mouse models and post-mortem brain and spinal cord tissues of ALS patients. Notably, MMR depletion resulted in the partial rescue of TDP43 proteinopathy-induced DNA damage and signaling. Moreover, bioinformatics analysis of the TCGA cancer database reveals significant associations between TDP43 expression, MMR gene expression, and mutational burden across multiple cancers. Collectively, our findings implicate TDP43 as a critical regulator of the MMR pathway and unveil its broad impact on the etiology of both neurodegenerative and neoplastic pathologies.

Amyotrophic lateral sclerosis↗