Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Computer automated prediction of potential therapeutic and toxicity protein targets of bioactive compounds from Chinese medicinal plants.

Understanding the molecular mechanism and pharmacology of bioactive compounds from Chinese medicinal plants (CMP) is important in facilitating scientific evaluation of novel therapeutic approaches in traditional Chinese medicine. It is also of significance in new drug development based on the mechanism of Chinese medicine. A key step towards this task is the determination of the therapeutic and toxicity protein targets of CMP compounds. In this work, newly developed computer software INVDOCK is used for automated identification of potential therapeutic and toxicity targets of several bioactive compounds isolated from Chinese medicinal plants. This software searches a protein database to find proteins to which a CMP compound can bind or weakly bind. INVDOCK results on three CMP compounds (allicin, catechin and camptotecin) show that 60% of computer-identified potential therapeutic protein targets and 27% of computer-identified potential toxicity targets have been implicated or confirmed by experiments. This software may potentially be used as a relatively fast-speed and low-cost tool for facilitating the study of molecular mechanism and pharmacology of bioactive compounds from Chinese medicinal plants and natural products from other sources.

Animals↗

Global genechip profiling to identify genes responsive to p53-induced growth arrest and apoptosis in human lung carcinoma cells.

To identify critical genes that mediate p53-induced growth arrest and apoptosis at a global level, we profiled a human lung carcinoma cell model in which cells undergo growth arrest and apoptosis in a p53 and DNA damage-dependent manner. Profiling of the Affymetrix human HG-U1333 GeneChip, covering the entire human transcriptome, revealed about 3, 000 unique genes either induced or repressed during p53-induced growth arrest or apoptosis, respectively. A total of 1, 057 genes, including many well-known p53 targets, responded to both conditions. A mini apoptotic protein database was generated from 3, 033 unique apoptosis responsive genes. Analysis of this database yielded 23 proteins with a pro-apoptotic BH3 domain and three with anti-apoptotic BIR2/BIR3 domains, including well-known p53 targets: Bax, Puma, Noxa and survivin. In addition, 14 mitochondrial proteins were identified that contain a pro-apoptotic AVPI-like motif, and 15 proteins were identified that contain a DAVPI-like domain with the potential of being cleaved by caspases during apoptosis to release the AVPI motif. Many of the genes we identified with these domains do contain p53-binding sites either in the promoter or in the first three introns, suggesting a high probability of being direct p53 targets. Pathway analysis revealed that p53 might control the Wnt pathway through transcriptional regulation of some of its components. Thus, global chip profiling coupled with bioinformatics analysis is a powerful tool in identification of genes critical for p53-induced apoptosis. Further characterization of these genes will lead to a better understanding of the mechanism of p53 action and p53 regulation of other signaling pathways. It will also provide novel cancer drug targets for further validation.

Apoptosis↗

[Protein profiling of human dendritic cells infected with mycobacterium tuberculosis].

The response of dendritic cells (DCs) plays an essential role in the initiation of immune responses following Mycobacterium tuberculosis (MTB) challenge. Two-dimensional electrophoresis (2-DE) was employed to compare the global protein patterns between human DCs infected and that uninfected with MTB H37Rv ATCC 27294 strains, and 45 protein spots were found to express differentially. Four protein spots which remarkably changed in DCs infected with MTB H37Rv ATCC 27294 strains were measured by matrix assisted laser desorption/ionization tandem time-of-flight (TOF/TOF) mass spectrometry. The data obtained from peptide mass fingerprinting were used in protein database search. Four protein spots in gel were identified as Human Arsenite-stimulated ATPase (hASNA-I), Annexin IV, gamma-actin and Heat shock protein27 (HSP27). These data provide insight into the changed global protein patterns of the DCs after infection and may prove useful for further study in the interaction between MTB and host.

Actins↗

[Preliminary proteomics analysis of the total proteins of HL Type cytoplasmic male sterility rice anther].

The proteins of HL type cytoplasmic male sterility rice anther of YTA (CMS) and YTB (maintenance line) were separated by two-dimensional electrophoresis with immobilized ph (3-10 non-linear) gradients as the first dimension and SDS-PAGE as the second. The silver-stained proteins spots were analyzed using Image Master 2D software, there were about 1800 detectable spots on each 2D-gel, and about 85 spots were differential expressed. With direct MALDI-TOF mass spectrometry analysis and protein database searching, 9 protein spots out of 16 were identified. Among those proteins, there were Putative nucleic acid binding protein, glucose-1-phosphate adenylyltransferase (ADP-glucose pyrophosphorylase, AGPase) (EC: 2.7.7.27) large chain, UDP-glucuronic acid decarboxylase, putative calcium-binding protein annexin, putative acetyl-CoA synthetase and putative lipoamide dehydrogenase etc. They were closely associated with metabolism, protein biosynthesis, transcription, signal transduction and so on, all of which are cell activities that are essential to pollen development. Some of the identified proteins, i.e. AGPase, putative lipoamide dehydrogenase and putative acetyl-CoA synthetase were deeply discussed on the relationship to CMS. AGPase catalyzes a very important step in the biosynthesis of alpha 1,4-glucans (glycogen or starch) in bacteria and plants: synthesis of the activated glucosyl donor, ADP-glucose, from glucose-1-phosphate and ATP. The lack of the AGPase in male sterile line might directly result in the reduction of starch, and the synthesis of starch was the most important processes during the development of pollen. In present research, the descent or reduction of putative lipoamide dehydrogenase and putative acetyl-CoA synthetase seemed involved in pollen sterility in rice. The degeneration and formation of various tissues during pollen development may impose high demands for energy and key biosynthetic intermediates. Under such conditions, the TCA cycle needs to operate fully, because the TCA cycle is an important source for many intermediates required for biosynthetic pathways, in addition to performing an oxidative, energy-producing role. Thus, it seemed reasonable to infer that the decrease of putative lipoamide dehydrogenase and putative acetyl-CoA synthetase in anther might prevent the conversion of pyruvate into acetyl-CoA, and as a result, the TCA cycle could no longer operate at a sufficient rate to meet all requirements in anther cells, leading to pollen sterility. This study gave new insights into the mechanism of CMS in rice and demonstrated the power of the proteomic approach in plant biology studies.

Electrophoresis, Polyacrylamide Gel↗

Interpreting peptide mass spectra by VEMS.

Most existing Mass Spectra (MS) analysis programs are automatic and provide limited opportunity for editing during the interpretation. Furthermore, they rely entirely on publicly available databases for interpretation. VEMS (Virtual Expert Mass Spectrometrist) is a program for interactive analysis of peptide MS/MS spectra imported in text file format. Peaks are annotated, the monoisotopic peaks retained, and the b-and y-ion series identified in an interactive manner. The called peptide sequence is searched against a local protein database for sequence identity and peptide mass. The report compares the calculated and the experimental mass spectrum of the called peptide. The program package includes four accessory programs. VEMStrans creates protein databases in FASTA format from EST or cDNA sequence files. VEMSdata creates a virtual peptide database from FASTA files. VEMSdist displays the distribution of masses up to 5000 Da. VEMSmaldi searches singly charged peptide masses against the local database.

Databases, Bibliographic↗

Identification of a key protein associated with cerebral ischemia.

The biochemical effects of permanent focal ischemia following unilateral occlusion of the middle cerebral artery in rats were studied by determining the content of specific proteins of the affected areas in the cerebral hemisphere. Brain proteins were prepared 72 h after the occlusion and analyzed by sodium dodecylsulfate-polyacrylamide gel electrophoresis. A significant increase in 66 and 80 kDa components and a paradoxical decrease in 260 kDa protein occurred in the ischemic brain tissues. The 66 and 80 kDa protein bands were identified as albumin and transferrin, respectively. The 260 kDa protein was analyzed by peptide mass fingerprinting (PMF) and matrix assisted laser desorption ionization time of flight mass spectrometry (MALDI-TOF-MS). The isoelectric point of the 260 kDa protein was 4.65 determined by isoelectric focusing. The data obtained from PMF were used in searching the protein database for homologous components. Three proteins with partial homology were identified. They were the microtubule-associated protein 1A, protein-tyrosine phosphatase zeta precursor (phosphacan), and protein kinase A anchoring protein 6. Polyclonal antibodies against the 260 kDa protein were raised and used to immunolocalize the antigen in various tissues. Positive staining occurred with brain neurons and pyramidal cells, islet cells, podocytes of kidney glomeruli, and endothelial cells of the venous sinuses of the spleen. The localization of 260 kDa protein strongly implies its function in these tissues. Its physiological and pathophysiological significances need to be clarified in future.

Animals↗

Capillary chromatography/microelectrospray mass spectrometry used for the identification of putative cyclin-dependent kinase inhibitory protein in Medicago.

A home-built capillary chromatography/microelectrospray system was used for peptide mass mapping of a putative cyclin-dependent kinase inhibitor protein with molecular mass of 8.5 kDa. The masses of identified tryptic fragments were then input to a protein database search routine. Daughter ion scans were done only on the most abundant tryptic fragments. On the basis of protein database search results and tandem mass spectrometric measurements the bioactive protein was identified as ubiquitin.

Amino Acid Sequence↗

FlexX-Scan: fast, structure-based virtual screening.

We present a new software module, FlexX-Scan, for high-throughput, structure-based virtual screening. FlexX-Scan was developed with the aim to further speed up the virtual screening process. Based on the incremental construction docking tool FlexX (Rarey et al., J Mol Biol 1996;261:470-489), a compact descriptor for representing favorable protein interaction spots within the protein binding site has been developed. The descriptor is calculated using special-purpose clustering techniques applied to the usual interaction points created by FlexX. The algorithm automatically detects a small set of interaction spots in the binding site for positioning ligand functional groups. The parametrizations of the base placement and incremental construction algorithms have been adapted to the new interaction model. We tested the software tool on a diverse set of 200 protein-ligand complexes from the protein database (PDB) (Kramer et al., Proteins 1999;37:228-241). On average, the algorithm proposes about 90 interaction spots per binding site compared to about 1000 interaction dots in FlexX. We observe that the docking solutions of FlexX-Scan have a root-mean-square deviation from the crystal structure similar to the deviation of docking solutions of standard FlexX. For further validation we also performed virtual screening experiments for cyclin-dependent kinase 2, thrombin, angiotensin-converting enzyme, and dihydrofolat reductase. In these experiments, we screened a set of 34,000 random compounds and a number of known actives for each target. With FlexX-Scan, we achieved comparable enrichments to standard FlexX, with an averaged computing time of 5-10 s per compound, depending on parametrization.

Algorithms↗

Tumorigenic poxviruses: genomic organization and DNA sequence of the telomeric region of the Shope fibroma virus genome.

Shope fibroma virus (SFV), a tumorigenic poxvirus, has a 160-kb linear double-stranded DNA genome and possesses terminal inverted repeats (TIRs) of 12.4 kb. The DNA sequence of the terminal 5.5 kb of the viral genome is presented and together with previously published sequences completes the entire sequence of the SFV TIR. The terminal 400-bp region contains no major open reading frames (ORFs) but does possess five related imperfect palindromes. The remaining 5.1 kb of the sequence contains seven tightly clustered and tandemly oriented ORFs, four larger than 100 amino acids in length (T1, T2, T4, and T5) and three smaller ORFs (T3A, T3B, and T3C). All are transcribed toward the viral hairpin and almost all possess the consensus sequence TTTTTNT near their 3' ends which has been implicated for the transcription termination of vaccinia virus early genes. Searches of the published DNA database revealed no sequences with significant homology with this region of the SFV genome but when the protein database was searched with the translation products of ORFs T1-T5 it was found that the N-terminus of the putative T4 polypeptide is closely related to the signal sequence of the hemagglutinin precursor from influenza A virus, suggesting that the T4 polypeptide may be secreted from SFV-infected cells. Examination of other SFV ORFs shows that T1 and T2 also possess signal-like hydrophobic amino acid stretches close to their N-termini. The protein database search also revealed that the putative T2 protein has significant homology to the insulin family of polypeptides. In terms of sequence repetitions, seven tandemly repeated copies of the hexanucleotide ATTGTT and three flanking regions of dyad symmetry were detected, all in ORF T3C. A search for palindromic sequences also revealed two clusters, one in ORF T3A/B and a second in ORF T2. ORF T2 harbors five short sequence domains, each of which consists of a 6-bp short palindrome and a 10- to 18-bp larger palindrome. The significance of these palindromic domains in this ORF is unclear but the coincidence of the end of one larger palindrome with the end of the translated protein sequence that has homology with the B chain of insulin suggests that the palindromes may divide the T2 protein into several functional units. The salient organizational features of the complete SFV TIR are also discussed in light of what is known about other poxviral TIRs.

Amino Acid Sequence↗

Comparison of methods for searching protein sequence databases.

We have compared commonly used sequence comparison algorithms, scoring matrices, and gap penalties using a method that identifies statistically significant differences in performance. Search sensitivity with either the Smith-Waterman algorithm or FASTA is significantly improved by using modern scoring matrices, such as BLOSUM45-55, and optimized gap penalties instead of the conventional PAM250 matrix. More dramatic improvement can be obtained by scaling similarity scores by the logarithm of the length of the library sequence (In()-scaling). With the best modern scoring matrix (BLOSUM55 or JO93) and optimal gap penalties (-12 for the first residue in the gap and -2 for additional residues), Smith-Waterman and FASTA performed significantly better than BLASTP. With In()-scaling and optimal scoring matrices (BLOSUM45 or Gonnet92) and gap penalties (-12, -1), the rigorous Smith-Waterman algorithm performs better than either BLASTP and FASTA, although with the Gonnet92 matrix the difference with FASTA was not significant. Ln()-scaling performed better than normalization based on other simple functions of library sequence length. Ln()-scaling also performed better than scores based on normalized variance, but the differences were not statistically significant for the BLOSUM50 and Gonnet92 matrices. Optimal scoring matrices and gap penalties are reported for Smith-Waterman and FASTA, using conventional or In()-scaled similarity scores. Searches with no penalty for gap extension, or no penalty for gap opening, or an infinite penalty for gaps performed significantly worse than the best methods. Differences in performance between FASTA and Smith-Waterman were not significant when partial query sequences were used. However, the best performance with complete query sequences was obtained with the Smith-Waterman algorithm and In()-scaling.

Algorithms↗

Mining protein data from two-dimensional gels: tools for systematic post-planned analyses.

There is a considerable need to develop comprehensive, systematic mechanisms to analyze the vast number of proteins that orchestrate various cellular functions and to identify proteins associated with disease or that are affected by pharmacological agents. Two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) continues to be relied upon to analyze protein constituents of cells and tissues. We have developed a Laboratory Information Processing System (LIPS) as a computer-based tool for capturing quantitative and qualitative changes in thousands of proteins detected in 2-D gels of various types. Protein databases have been developed to serve as a repository for data processing of the basic and derived data and of findings derived from different studies. There have been remarkable advances both in database technology as well as in the computer hardware that have benefited our effort at mining protein data from 2-D gels. We here review our current efforts aimed at improving the performance and features of our 2-D related protein databases, with particular emphasis on the tools we utilize for database mining via a systematic analysis of information known as post-planned analysis.

Acrylic Resins↗

RT-PSM, a real-time program for peptide-spectrum matching with statistical significance.

The analysis of complex biological peptide mixtures by tandem mass spectrometry (MS/MS) produces a huge body of collision-induced dissociation (CID) MS/MS spectra. Several methods have been developed for identifying peptide-spectrum matches (PSMs) by assigning MS/MS spectra to peptides in a database. However, most of these methods either do not give the statistical significance of PSMs (e.g., SEQUEST) or employ time-consuming computational methods to estimate the statistical significance (e.g., PeptideProphet). In this paper, we describe a new algorithm, RT-PSM, which can be used to identify PSMs and estimate their accuracy statistically in real time. RT-PSM first computes PSM scores between an MS/MS spectrum and a set of candidate peptides whose masses are within a preset tolerance of the MS/MS precursor ion mass. Then the computed PSM scores of all candidate peptides are employed to fit the expectation value distribution of the scores into a second-degree polynomial function in PSM score. The statistical significance of the best PSM is estimated by extrapolating the fitting polynomial function to the best PSM score. RT-PSM was tested on two pairs of MS/MS spectrum datasets and protein databases to investigate its performance. The MS/MS spectra were acquired using an ion trap mass spectrometer equipped with a nano-electrospray ionization source. The results show that RT-PSM has good sensitivity and specificity. Using a 55,577-entry protein database and running on a standard Pentium-4, 2.8-GHz CPU personal computer, RT-PSM can process peptide spectra on a sequential, one-by-one basis in 0.047 s on average, compared to more than 7 s per spectrum on average for Sequest and X!Tandem, in their current batch-mode processing implementations. RT-PSM is clearly shown to be fast enough for real-time PSM assignment of MS/MS spectra generated every 3 s or so by a 3D ion trap or by a QqTOF instrument.

Algorithms↗

MvirDB--a microbial database of protein toxins, virulence factors and antibiotic resistance genes for bio-defence applications.

Knowledge of toxins, virulence factors and antibiotic resistance genes is essential for bio-defense applications aimed at identifying 'functional' signatures for characterizing emerging or engineered pathogens. Whereas genetic signatures identify a pathogen, functional signatures identify what a pathogen is capable of. To facilitate rapid identification of sequences and characterization of genes for signature discovery, we have collected all publicly available (as of this writing), organized sequences representing known toxins, virulence factors, and antibiotic resistance genes in one convenient database, which we believe will be of use to the bio-defense research community. MvirDB integrates DNA and protein sequence information from Tox-Prot, SCORPION, the PRINTS virulence factors, VFDB, TVFac, Islander, ARGO and a subset of VIDA. Entries in MvirDB are hyperlinked back to their original sources. A blast tool allows the user to blast against all DNA or protein sequences in MvirDB, and a browser tool allows the user to search the database to retrieve virulence factor descriptions, sequences, and classifications, and to download sequences of interest. MvirDB has an automated weekly update mechanism. Each protein sequence in MvirDB is annotated using our fully automated protein annotation system and is linked to that system's browser tool. MvirDB can be accessed at http://mvirdb.llnl.gov/.

Bacterial Proteins↗

Analysis of the HUPO Brain Proteome reference samples using 2-D DIGE and 2-D LC-MS/MS.

Within the Human Proteome Organization (HUPO) Brain Proteome Project, a pilot study was launched with reference samples shipped to nine international laboratories (see Hamacher et al., this Special Issue) to evaluate different proteome approaches in neuroscience and to build up a first version of a brain protein database. One part of the study addresses quantitative proteome alterations between three developmental stages (embryonic day 16; postnatal day 7; 8 weeks) of mouse brains. Five brains per stage were differentially analyzed by 2-D DIGE using internal standardization and overlapping pH gradients (pH 4-7 and 6-9). In total, 214 protein spots showing stage-dependent intensity alterations (> two-fold) were detected, 56 of which were identified. Several of them, e.g. members of the dihydropyrimidinase family, are known to be associated with brain development. To feed the HUPO BPP brain protein database, a robust 2-D LC-MS/MS method was applied to murine postnatal day 7 and human post-mortem brain samples. Using MASCOT and the IPI database, 350 human and 481 mouse proteins could be identified by at least two different peptides. The data are accessible through the PRIDE database (http://www.ebi.ac.uk/pride/).

Aging↗

Associative database of protein sequences.

MOTIVATION: We present a new concept that combines data storage and data analysis in genome research, based on an associative network memory. As an illustration, 115 000 conserved regions from over 73 000 published sequences (i.e. from the entire annotated part of the SWISSPROT sequence database) were identified and clustered by a self-organizing network. Similarity and kinship, as well as degree of distance between the conserved protein segments, are visualized as neighborhood relationship on a two-dimensional topographical map. RESULTS: Such a display overcomes the restrictions of linear list processing and allows local and global sequence relationships to be studied visually. Families are memorized as prototype vectors of conserved regions. On a massive parallel machine, clustering and updating of the database take only a few seconds; a rapid analysis of incoming data such as protein sequences or ESTs is carried out on present-day workstations. AVAILABILITY: Access to the database is available at http://www.bioinf.mdc-berlin.de/unter2.html++ + CONTACT: (hanke,lehmann,reich)@mdc-berlin.de; bork@embl-heidelberg.de

Amino Acid Sequence↗