Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

ProtoBee: hierarchical classification and annotation of the honey bee proteome.

The recently sequenced genome of the honey bee (Apis mellifera) has produced 10,157 predicted protein sequences, calling for a computational effort to extract biological insights from them. We have applied an unsupervised hierarchical protein-clustering method, which was previously used in the ProtoNet system, to nearly 200,000 proteins consisting of the predicted honey bee proteins, the SWISS-PROT protein database, and the complete set of proteins of the mouse (Mus musculus) and the fruit fly (Drosophila melanogaster). The hierarchy produced by this method has been entitled ProtoBee. In ProtoBee, the proteins are hierarchically organized into 18,936 separate tree hierarchies, each representing a protein functional family. By using the mouse and Drosophila complete proteomes as reference, we are able to highlight functional groups of putative gene-loss events, putative novel proteins of unique functionality, and bee-specific paralogs. We have studied some of the ProtoBee findings and suggest their biological relevance. Examples include novel opsin genes and intriguing nuclear matches of mitochondrial genes. The organization of bee sequences into functional clusters suggests a natural way of automatically inferring functional annotation. Following this notion, we were able to assign functional annotation to about 70% of the sequences. ProtoBee is available at http://www.protobee.cs.huji.ac.il.

Animals↗

Sequencing of three lambda clones from the genome of alkaliphilic Bacillus sp. strain C-125.

The nucleotide sequences of three independent fragments (designated no. 3, 4, and 9; each 15-20 kb in size) of the genome of alkaliphilic Bacillus sp. C-125 cloned in a lambda phage vector have been determined. Thirteen putative open reading frames (ORFs) were identified in sequenced fragment no. 3 and 11 ORFs were identified in no. 4. Twenty ORFs were also identified in fragment no. 9. All putative ORFs were analyzed in comparison with the BSORF database and non-redundant protein databases. The functions of 5 ORFs in fragment no. 3 and 3 ORFs in fragment no. 4 were suggested by their significant similarities to known proteins in the database. Among the 20 ORFs in fragment no. 9, the functions of 11 ORFs were similarly suggested. Most of the annotated ORFs in the DNA fragments of the genome of alkaliphilic Bacillus sp. C-125 were conserved in the Bacillus subtilis genome. The organization of ORFs in the genome of strain C-125 was found to differ from the order of genes in the chromosome of B. subtilis, although some gene clusters (ydh, yqi, yer, and yts) were conserved as operon units the same as in B. subtilis.

Bacillus↗

Uncovering conserved patterns in bioactive peptides in Metazoa.

Bioactive (neuro)peptides play critical roles in regulating most biological processes in animals. Peptides belonging to the same family are characterized by a typical sequence pattern that is conserved among the family's peptide members. Such a conserved pattern or motif usually corresponds to the functionally important part of the biologically active peptide. In this paper, all known bioactive (neuro)peptides annotated in Swiss-Prot and TrEMBL protein databases are collected, and the pattern searching program Pratt is used to search these unaligned peptide sequences for conserved patterns. The obtained patterns are then refined by combining the information on amino acids at important functional sites collected from the literature. All the identified patterns are further tested by scanning them against Swiss-Prot and TrEMBL protein databases. The diagnostic power of each pattern is validated by the fact that any annotated protein from Swiss-Prot and TrEMBL that contains one of the established patterns, is indeed a known (neuro)peptide precursor. We discovered 155 novel peptide patterns in addition to the 56 established ones in the PROSITE database. All the patterns cover 110 peptide families. Fifty-five of these families are not characterized by the PROSITE signatures, and 12 are also not identified by other existing motif databases, such as Pfam and SMART. Using the newly identified peptide signatures as a search tool, we predicted 95 hypothetical proteins as putative peptide precursors.

Amino Acid Motifs↗

Convergence of amino acid compositions of certain groups of protein aids in their identification on two-dimensional electrophoresis gels.

The amino acid composition (AAC) versus the protein identity (PI) method was used for establishment of the identities of proteins from bovine brain and kidneys which were prefractionated on a CM52 cation exchanger and by preparative flat-bed isoelectric focusing. Established identities of proteins whose AACs converge with those of other members of their proper superfamily are reliable. Groups of convergent AACs can be extracted from protein databases using the standard root-mean-square rule (Rmsd) with measures the difference between the AAC of chosen protein versus those in the database. Convergence of AACs of proteins is dependent on several factors such as the upper limit of Rmsd, the limits of variations of molecular mass (m) and isoelectric point (pI), the number of proteins with similar AACs present in protein databases, and the domain structure of proteins. AACs of many proteins remain unique if the Rmsd is maintained within 1.5-1.0 with m +/- 3kDA and pI +/- 4. Certain groups of multidomain proteins have quasi-unique AACs only if the Rmsd is restrained to a value within 1.0 and 0.7. Convergence of AACs of certain groups of proteins may indicate that a common biological function exists for some members of each group. The AAC-PI method may become an additional search tool for protein functions.

Amino Acid Isomerases↗

PROSITE: a documented database using patterns and profiles as motif descriptors.

Among the various databases dedicated to the identification of protein families and domains, PROSITE is the first one created and has continuously evolved since. PROSITE currently consists of a large collection of biologically meaningful motifs that are described as patterns or profiles, and linked to documentation briefly describing the protein family or domain they are designed to detect. The close relationship of PROSITE with the SWISS-PROT protein database allows the evaluation of the sensitivity and specificity of the PROSITE motifs and their periodic reviewing. In return, PROSITE is used to help annotate SWISS-PROT entries. The main characteristics and the techniques of family and domain identification used by PROSITE are reviewed in this paper.

Amino Acid Motifs↗

Coiled-coil protein composition of 22 proteomes--differences and common themes in subcellular infrastructure and traffic control.

BACKGROUND: Long alpha-helical coiled-coil proteins are involved in diverse organizational and regulatory processes in eukaryotic cells. They provide cables and networks in the cyto- and nucleoskeleton, molecular scaffolds that organize membrane systems and tissues, motors, levers, rotating arms, and possibly springs. Mutations in long coiled-coil proteins have been implemented in a growing number of human diseases. Using the coiled-coil prediction program MultiCoil, we have previously identified all long coiled-coil proteins from the model plant Arabidopsis thaliana and have established a searchable Arabidopsis coiled-coil protein database. RESULTS: Here, we have identified all proteins with long coiled-coil domains from 21 additional fully sequenced genomes. Because regions predicted to form coiled-coils interfere with sequence homology determination, we have developed a sequence comparison and clustering strategy based on masking predicted coiled-coil domains. Comparing and grouping all long coiled-coil proteins from 22 genomes, the kingdom-specificity of coiled-coil protein families was determined. At the same time, a number of proteins with unknown function could be grouped with already characterized proteins from other organisms. CONCLUSION: MultiCoil predicts proteins with extended coiled-coil domains (more than 250 amino acids) to be largely absent from bacterial genomes, but present in archaea and eukaryotes. The structural maintenance of chromosomes proteins and their relatives are the only long coiled-coil protein family clearly conserved throughout all kingdoms, indicating their ancient nature. Motor proteins, membrane tethering and vesicle transport proteins are the dominant eukaryote-specific long coiled-coil proteins, suggesting that coiled-coil proteins have gained functions in the increasingly complex processes of subcellular infrastructure maintenance and trafficking control of the eukaryotic cell.

Amino Acid Sequence↗

Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.

BACKGROUND: In order to improve gene prediction, extrinsic evidence on the gene structure can be collected from various sources of information such as genome-genome comparisons and EST and protein alignments. However, such evidence is often incomplete and usually uncertain. The extrinsic evidence is usually not sufficient to recover the complete gene structure of all genes completely and the available evidence is often unreliable. Therefore extrinsic evidence is most valuable when it is balanced with sequence-intrinsic evidence. RESULTS: We present a fairly general method for integration of external information. Our method is based on the evaluation of hints to potentially protein-coding regions by means of a Generalized Hidden Markov Model (GHMM) that takes both intrinsic and extrinsic information into account. We used this method to extend the ab initio gene prediction program AUGUSTUS to a versatile tool that we call AUGUSTUS+. In this study, we focus on hints derived from matches to an EST or protein database, but our approach can be used to include arbitrary user-defined hints. Our method is only moderately effected by the length of a database match. Further, it exploits the information that can be derived from the absence of such matches. As a special case, AUGUSTUS+ can predict genes under user-defined constraints, e.g. if the positions of certain exons are known. With hints from EST and protein databases, our new approach was able to predict 89% of the exons in human chromosome 22 correctly. CONCLUSION: Sensitive probabilistic modeling of extrinsic evidence such as sequence database matches can increase gene prediction accuracy. When a match of a sequence interval to an EST or protein sequence is used it should be treated as compound information rather than as information about individual positions.

Algorithms↗

Protein composition of Paracoccus denitrificans cells grown on various electron acceptors and in the presence of azide.

Two-dimensional gel electrophoresis (2-DE) with immobilized pH gradients was carried out on total cell lysates and membrane fractions of Paracoccus denitrificans with the aim to characterize differences in protein expression during growth under aerobic and various anaerobic conditions (with nitrate, nitrite or nitrous oxide). Comparative image analysis of the protein pattern revealed several subgroups of the total 800 protein spots resolved that were characteristically induced or repressed in response to individual electron acceptors. The respiratory inhibitor azide also exerted a profound influence upon cellular protein composition. However, since most of the proteins showing an altered expression pattern in cells growing on oxygen differed from those in cells growing on nitrite, we suppose that azide acts mainly indirectly, possibly by influencing other cellular signals. Limited information on the P. denitrificans genome has precluded the identification of more than eight protein spots as yet. A public accessible P. denitrificans 2-DE protein database is currently built up at http://www.mpiib-berlin.mpg.de/2D-PAGE.

Bacterial Proteins↗

Modular, scriptable and automated analysis tools for high-throughput peptide mass fingerprinting.

UNLABELLED: A set of new algorithms and software tools for automatic protein identification using peptide mass fingerprinting is presented. The software is automatic, fast and modular to suit different laboratory needs, and it can be operated either via a Java user interface or called from within scripts. The software modules do peak extraction, peak filtering and protein database matching, and communicate via XML. Individual modules can therefore easily be replaced with other software if desired, and all intermediate results are available to the user. The algorithms are designed to operate without human intervention and contain several novel approaches. The performance and capabilities of the software is illustrated on spectra from different mass spectrometer manufacturers, and the factors influencing successful identification are discussed and quantified. MOTIVATION: Protein identification with mass spectrometric methods is a key step in modern proteomics studies. Some tools are available today for doing different steps in the analysis. Only a few commercial systems integrate all the steps in the analysis, often for only one vendor's hardware, and the details of these systems are not public. RESULTS: A complete system for doing protein identification with peptide mass fingerprints is presented, including everything from peak picking to matching the database protein. The details of the different algorithms are disclosed so that academic researchers can have full control of their tools. AVAILABILITY: The described software tools are available from the Halmstad University website www.hh.se/staff/bioinf/ SUPPLEMENTARY INFORMATION: Details of the algorithms are described in supporting information available from the Halmstad University website www.hh.se/staff/bioinf/

Algorithms↗

Side-chain interactions between sulfur-containing amino acids and phenylalanine in alpha-helices.

The side-chain-side-chain interaction between Phe residues and sulfur-containing residues (Cis and Met) in the two possible orientations at positions i, i + 4 of alpha-helices is described. We have analyzed the contribution to helical stability of the above interactions by studying eight polyalanine-based peptides differing at the residues at positions 9 and 13. These two positions were independently mutated from Ala (AA), to Cys (AC and CA), Met (AM and MA), and Phe (AF and FA) and to the pairs Phe-Met (FM), Met-Phe (MF), Phe-Cys (FC), and Cys-Phe (CF). The intrinsic helical propensities of Cys, Met, and Phe were found to be those previously described in the algorithm AGADIR. NMR analysis of the FM, MF, FC, and CF peptides showed the formation in aqueous solution of contacts between the aromatic ring and the side chains of Cys or Met, at the two i, i + 4 orientations. CD studies demonstrated the important contribution of two of these interactions (FM and FC) to alpha-helix stability (up to 2 kcal mol-1 in the Phe-Cys pair). Statistical analysis of the protein database provides a rationale for the stereospecificity and free energies of the interactions. The very favorable interaction between an aromatic ring and a sulfur-containing amino acid explains why in the protein database around 50% of the sulfur atoms are contacting aromatic rings (Reid et al., 1985).

Amino Acid Sequence↗

The CATH database: an extended protein family resource for structural and functional genomics.

The CATH database of protein domain structures (http://www.biochem.ucl.ac.uk/bsm/cath_new) currently contains 34 287 domain structures classified into 1383 superfamilies and 3285 sequence families. Each structural family is expanded with domain sequence relatives recruited from GenBank using a variety of efficient sequence search protocols and reliable thresholds. This extended resource, known as the CATH-protein family database (CATH-PFDB) contains a total of 310 000 domain sequences classified into 26 812 sequence families. New sequence search protocols have been designed, based on these intermediate sequence libraries, to allow more regular updating of the classification. Further developments include the adaptation of a recently developed method for rapid structure comparison, based on secondary structure matching, for domain boundary assignment. The philosophy behind CATHEDRAL is the recognition of recurrent folds already classified in CATH. Benchmarking of CATHEDRAL, using manually validated domain assignments, demonstrated that 43% of domains boundaries could be completely automatically assigned. This is an improvement on a previous consensus approach for which only 10-20% of domains could be reliably processed in a completely automated fashion. Since domain boundary assignment is a significant bottleneck in the classification of new structures, CATHEDRAL will also help to increase the frequency of CATH updates.

Animals↗

Arabidopsis thaliana proteomics: from proteome to genome.

Proteomics has become an important approach for investigating cellular processes and network functions. Significant improvements have been made during the last few years in technologies for high-throughput proteomics, both at the level of data analysis software and mass spectrometry hardware. As proteomics technologies advance and become more widely accessible, efforts of cataloguing and quantifying full proteomes are underway to complement other genomics approaches, such as RNA and metabolite profiling. Of particular interest is the application of proteome data to improve genome annotation and to include information on post-translational protein modifications with the annotation of the corresponding gene. This type of analysis requires a paradigm shift because amino acid sequences must be assigned to peptides without relying on existing protein databases. In this review, advances and current limitations of full proteome analysis are briefly highlighted using the model plant Arabidopsis thaliana as an example. Strategies to identify peptides are also discussed on the basis of MS/MS data in a protein database-independent approach.

Arabidopsis↗

De novo identification of cell-type specific antibody-antigen pairs by phage display subtraction. Isolation of a human single chain antibody fragment against human keratin 14.

The aim of this study was to identify novel antibodies directed against cytosolic keratinocyte-specific antigens from a phage display antibody repertoire by using phage display subtraction. Phage display is a method of displaying foreign molecules on the surface of filamentous bacteriophage particles. It allows the interaction between two cognate molecules to be analysed through affinity selections. Recently, large repertoires of phage displayed human antibody fragments have been constructed. From such repertoires, antibodies can be obtained in vitro without the need for immunization or the hybridoma technology. A novel subtractive strategy for selecting antibodies from phage libraries was applied. Phage antibodies were selected against immobilized crude lysates of cultured human keratinocytes, the target antigens being unknown beforehand. A competing cell lysate was used to reduce retrieval of phage antibodies with specificities to commonly non-differentially expressed antigens. A monoclonal single chain fragment variable (scFv) with specificity for crude lysates of cultured human keratinocytes was identified as demonstrated by ELISA assays and immunoblotting analysis. The cognate keratinocyte antigen was shown to be keratin 14 (K14) by using immunoblotting based on 2D PAGE and a corresponding 2D PAGE protein database. In accordance with the expected tissue localization of K14, the identified scFv stained the basal layer of human epidermis by indirect immunofluorescence analysis. Starting with crude cell lysates, phage display subtraction in combination with 2D PAGE and 2D PAGE protein databases can be used to identify antibody-antigen pairs that characterize a specific cell type.

Blotting, Western↗

Predicting protein function: a versatile tool for the Apple Macintosh.

A tool is presented that helps to find biological functions for new protein sequences. Running on any Macintosh computer system, MacPattern provides a unique combination of different algorithms, speed and user-friendliness. It supports searches for protein patterns using the PROSITE database, protein block searches with the BLOCKS database, and the identification of statistically significant protein segments. MacPattern allows batch processing of sequences and automatic translations of nucleotide sequence data. It is particularly suited for genome analysis or cDNA sequencing projects.

Algorithms↗