Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Signature”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Untraceable partially blind signature based on DLOG problem.

This paper proposes a new untraceable Partially Blind Signature scheme which is a cross between the traditional signature scheme and the blind signature scheme. In this proposed scheme, the message M that the signer signed can be divided into two parts. The first part can be known to the signer (like that in the traditional signature scheme) while the other part cannot be known to the signer (like that in the blind signature scheme). After having signed M, the signer cannot determine if he has made the signature of M except through the part that he knows. We draw ideas from Brands' "Restricted Blind Signature" to solve the Untraceable Partially Blind Signature problem. Our scheme is a probabilistic signature scheme and the security of our Untraceable Partially Blind Signature scheme relies on the difficulty of computing discrete logarithm.

Algorithms↗

Discover protein sequence signatures from protein-protein interaction data.

BACKGROUND: The development of high-throughput technologies such as yeast two-hybrid systems and mass spectrometry technologies has made it possible to generate large protein-protein interaction (PPI) datasets. Mining these datasets for underlying biological knowledge has, however, remained a challenge. RESULTS: A total of 3108 sequence signatures were found, each of which was shared by a set of guest proteins interacting with one of 944 host proteins in Saccharomyces cerevisiae genome. Approximately 94% of these sequence signatures matched entries in InterPro member databases. We identified 84 distinct sequence signatures from the remaining 172 unknown signatures. The signature sharing information was then applied in predicting sub-cellular localization of yeast proteins and the novel signatures were used in identifying possible interacting sites. CONCLUSION: We reported a method of PPI data mining that facilitated the discovery of novel sequence signatures using a large PPI dataset from S. cerevisiae genome as input. The fact that 94% of discovered signatures were known validated the ability of the approach to identify large numbers of signatures from PPI data. The significance of these discovered signatures was demonstrated by their application in predicting sub-cellular localizations and identifying potential interaction binding sites of yeast proteins.

Binding Sites↗

Global transcriptional profiling of the toxic dinoflagellate Alexandrium fundyense using Massively Parallel Signature Sequencing.

BACKGROUND: Dinoflagellates are one of the most important classes of marine and freshwater algae, notable both for their functional diversity and ecological significance. They occur naturally as free-living cells, as endosymbionts of marine invertebrates and are well known for their involvement in "red tides". Dinoflagellates are also notable for their unusual genome content and structure, which suggests that the organization and regulation of dinoflagellate genes may be very different from that of most eukaryotes. To investigate the content and regulation of the dinoflagellate genome, we performed a global analysis of the transcriptome of the toxic dinoflagellate Alexandrium fundyense under nitrate- and phosphate-limited conditions using Massively Parallel Signature Sequencing (MPSS). RESULTS: Data from the two MPSS libraries showed that the number of unique signatures found in A. fundyense cells is similar to that of humans and Arabidopsis thaliana, two eukaryotes that have been extensively analyzed using this method. The general distribution, abundance and expression patterns of the A. fundyense signatures were also quite similar to other eukaryotes, and at least 10% of the A. fundyense signatures were differentially expressed between the two conditions. RACE amplification and sequencing of a subset of signatures showed that multiple signatures arose from sequence variants of a single gene. Single signatures also mapped to different sequence variants of the same gene. CONCLUSION: The MPSS data presented here provide a quantitative view of the transcriptome and its regulation in these unusual single-celled eukaryotes. The observed signature abundance and distribution in Alexandrium is similar to that of other eukaryotes that have been analyzed using MPSS. Results of signature mapping via RACE indicate that many signatures result from sequence variants of individual genes. These data add to the growing body of evidence for widespread gene duplication in dinoflagellates, which would contribute to the transcriptional complexity of these organisms. The MPSS data also demonstrate that a significant number of dinoflagellate mRNAs are transcriptionally regulated, indicating that dinoflagellates commonly employ transcriptional gene regulation along with the post-transcriptional regulation that has been well documented in these organisms.

Animals↗

Forensic handwriting examiners' expertise for signature comparison.

This paper reports on the performance of forensic document examiners (FDEs) in a signature comparison task that was designed to address the issue of expertise. The opinions of FDEs regarding 150 genuine and simulated questioned signatures were compared with a control group of non-examiners' opinions. On the question of expertise, results showed that FDEs were statistically better than the control group at accurately determining the genuineness or non-genuineness of questioned signatures. The FDE group made errors (by calling a genuine signature simulated or by calling a simulated signature genuine) in 3.4% of their opinions while 19.3% of the control group's opinions were erroneous. The FDE group gave significantly more inconclusive opinions than the control group. Analysis of FDEs' responses showed that more correct opinions were expressed regarding simulated signatures and more inconclusive opinions were made on genuine signatures. Further, when the complexity of a signature was taken into account, FDEs made more correct opinions on high complexity signatures than on signatures of lower complexity. There was a wide range of skill amongst FDEs and no significant relationship was found between the number of years FDEs had been practicing and their correct, inconclusive and error rates.

Adult↗

NIST bullet signature measurement system for RM (Reference Material) 8240 standard bullets.

A bullet signature measurement system based on a stylus instrument was developed at the National Institute of Standards and Technology (NIST) for the signature measurements of NIST RM (Reference Material) 8240 standard bullets. The standard bullets are developed as a reference standard for bullet signature measurements and are aimed to support the recently established National Integrated Ballistics Information Network (NIBIN) by the Bureau of Alcohol, Tobacco and Firearms (ATF) and the Federal Bureau of Investigation (FBI). The RM bullets are designed as both a virtual and a physical bullet signature standard. The virtual standard is a set of six digitized bullet signatures originally profiled from six master bullets fired at ATF and FBI using six different guns. By using the virtual signature standard to control the tool path on a numerically controlled diamond turning machine at NIST, 40 RM bullets were produced. In this paper, a comparison parameter and an algorithm using auto-and cross-correlation functions are described for qualifying the bullet signature differences between the RM bullets and the virtual bullet signature standard. When two compared signatures are exactly the same (point by point), their cross-correlation function (CCF) value will be equal to 100%. The measurement system setup, measurement program, and initial measurement results are discussed. Initial measurement results for the 40 standard bullets, each measured at six land impressions, show that the CCF values for the 240 signature measurements are higher than 95%, with most of them even higher than 99%. These results demonstrate the high reproducibility for both the manufacturing process and the measurement system for the NIST RM 8240 standard bullets.

Algorithms↗

[Identification of signature amino acids in the env V3-V4 and flanking regions of human immunodeficiency virus type 1 predominant strains in China].

OBJECTIVE: To identify signature amino acids in the V3-V4 and flanking regions of the env gene from human immunodeficiency virus type 1 (HIV-1) predominant strains in China and to elucidate the role of these signature amino acids on epidemiologic tracking and the development of vaccine. METHODS: Fragments of the HIV-1 env gene were amplified by nested-PCR from the whole blood of HIV-1 infected individuals from 12 provinces in China. Then, the PCR products were directly sequenced by using ABI 377 DNA SEQUENCER. The sequences covering the env V3-V4 region of the strains were used for the analyses described here. Envelope sequence subtypes were assigned using BLAST (http://www.HIV-Web.lanl.gov). Phylogenetic analyses were performed using GCG and MEGA as well as signature amino acids were identified using VESPA. RESULTS: Subtype B' strains and two recombinants (B'/C and CRF01-AE) were discovered among 157 currently circulating strains in China. The most prevalent subtypes were B'/C (38.85%), followed by B' (34.40%), and CRF01-AE (26.75%). Phylogenetic tree analysis of env V3-V4 region showed that subtype B' strains were closely related to B.CN.RL42, while most of B'/C strains clustered with 97CN54A and 97CNGX6F, CRF01-AE strains clustered into two distinct subgroups, which were closely related to THCM240 and 97CNGX2F. Analysis of signature amino acids revealed that eight of positions were identified as conserved signature amino acid sites in the env V3-V4 and flanking regions of the subtype B' and B'/C strains and almost all signature amino acids were found in their reference strains. Interestingly, eleven signature amino acids were demonstrated in the same regions of the CRF01-AE strains, but nine out of 11 signature amino acid sites were distinct from the same positions of the reference strains 97CNGX2F and TH.CM240. It is noteworthy that these 9 signature amino acids were found in the strains from all of the selected provinces except those from Yunnan province. CONCLUSION: Analysis of the signature amino acids suggest that much of the current Chinese epidemics of subtype B' and B'/C strains are descended from a single introduction into China, while the epidemic caused by the CRF01-AE strain is caused by multiple introductions into China from Thailand. These results will contribute to the policy of AIDS prevention and control as well as the ongoing development of AIDS vaccine.

Adolescent↗

Quantifying the species-specificity in genomic signatures, synonymous codon choice, amino acid usage and G+C content.

Each prokaryote has a unique genomic signature as evidenced by a set of species-specific frequencies of short oligonucleotides. With respect to genomic signatures a bacterial genome is homogenous and the variation within a genome is smaller than the variations between genomes of different species. This study quantifies the species-specificity of genomic signatures in the complete genomes of 57 prokaryotes. The species-specificity in the genomic signature was related to the quantification of other sequence biases, such as G+C content, synonymous codon choice and amino acid usage. The results confirm that the genomic signature is genome-wide with high species-specificity in both coding and non-coding regions. In coding regions the species-specific bias in synonymous codon choice was comparable to the genomic signature, while the bias in amino acid usage only captured about 50% of the species-specific bias in the genomic signature. A correlation between the species-specificity in synonymous codon choice and amino acid usage was identified, in which proteins with species-specific amino acid usage were also coded with species-specific synonymous codon choice. However, we demonstrated that the G+C content captures only approximately 40% of the species-specificity in the genomic signature, and is insufficient to explain the species specificity in the non-coding regions. Thus, the species-specific bias in non-coding regions remains largely unknown. Further, we compared the genomic signature in relation to phylogenetic distance. This was performed in order to illustrate the feasibility of a hierarchical classification scheme in future applications of the described classification methodology in screening for horizontal gene transfer and biodiversity studies.

Amino Acids↗

The signature molecular descriptor. 1. Using extended valence sequences in QSAR and QSPR studies.

We present a new descriptor named signature based on extended valence sequence. The signature of an atom is a canonical representation of the atom's environment up to a predefined height h. The signature of a molecule is a vector of occurrence numbers of atomic signatures. Two QSAR and QSPR models based on signature are compared with models obtained using popular molecular 2D descriptors taken from a commercially available software (Molconn-Z). One set contains the inhibition concentration at 50% for 121 HIV-1 protease inhibitors, while the second set contains 12865 octanol/water partitioning coefficients (Log P). For both data sets, the models created by signature performed comparable to those from the commercially available descriptors in both correlating the data and in predicting test set values not used in the parametrization. While probing signature's QSAR and QSPR performances, we demonstrates that for any given molecule of diameter D, there is a molecular signature of height h </= D+1, from which any 2D descriptor can be computed. As a consequence of this finding any QSAR or QSPR involving 2D descriptors can be replaced with a relationship involving occurrence number of atomic signatures.

Algorithms↗

The signature molecular descriptor. 2. Enumerating molecules from their extended valence sequences.

We present a new algorithm that enumerates molecular structures matching a predefined extended valence sequence or signature. The algorithm can construct molecular structures composed of about 50 non-hydrogen atoms in CPU seconds time scale. The algorithm is run to produce all molecular structures matching the binding affinities (IC(50)) of some HIV-1 protease inhibitors. The algorithm is also used to compute the degeneracy, or the number of molecular structures, corresponding to a given signature. Signature degeneracy is systematically studied for varying signature heights on four molecular series, alkanes, alcohols, fullerene-type structures, and peptides. Signature degeneracy is compared with similar results obtained with popular topological indices (TIs). As a general rule, we find that signature degeneracy decreases as the signature height increases. We also find that alkanes, alcohols, and fullerene-type structures comprising n non-hydrogen atoms are uniquely characterized by signatures of height n/4, while peptides up to 4000 amino acids can be singled out with signatures of heights as small as 2 and 3.

Algorithms↗

Protein signatures distinctive of alpha proteobacteria and its subgroups and a model for alpha-proteobacterial evolution.

Alpha (alpha) proteobacteria comprise a large and metabolically diverse group. No biochemical or molecular feature is presently known that can distinguish these bacteria from other groups. The evolutionary relationships among this group, which includes numerous pathogens and agriculturally important microbes, are also not understood. Shared conserved inserts and deletions (i.e., indels or signatures) in molecular sequences provide a powerful means for identification of different groups in clear terms, and for evolutionary studies (see www.bacterialphylogeny.com). This review describes, for the first time, a large number of conserved indels in broadly distributed proteins that are distinctive and unifying characteristics of either all alpha-proteobacteria, or many of its constituent subgroups (i.e., orders, families, etc.). These signatures were identified by systematic analyses of proteins found in the Rickettsia prowazekii (RP) genome. Conserved indels that are unique to alpha-proteobacteria are present in the following proteins: Cytochrome c oxidase assembly protein Ctag, PurC, DnaB, ATP synthase alpha-subunit, exonuclease VII, prolipoprotein phosphatidylglycerol transferase, RP-400, FtsK, puruvate phosphate dikinase, cytochrome b, MutY, and homoserine dehydrogenase. The signatures in succinyl-CoA synthetase, cytochrome oxidase I, alanyl-tRNA synthetase, and MutS proteins are found in all alpha-proteobacteria, except the Rickettsiales, indicating that this group has diverged prior to the introduction of these signatures. A number of proteins contain conserved indels that are specific for Rickettsiales (XerD integrase and leucine aminopeptidase), Rickettsiaceae (Mfd, ribosomal protein L19, FtsZ, Sigma 70 and exonuclease VII), or Anaplasmataceae (Tgt and RP-314), and they distinguish these groups from all others. Signatures in DnaA, RP-057, and DNA ligase A are commonly shared by various Rhizobiales, Rhodobacterales, and Caulobacter, suggesting that these groups shared a common ancestor exclusive of other alpha-proteobacteria. A specific relationship between Rhodobacterales and Caulobacter is indicated by a large insert in the Asn-Gln amidotransferase. The Rhizobiales group of species are distinguished from others by a large insert in the Trp-tRNA synthetase. Signature sequences in a number of other proteins (viz. oxoglutarate dehydogenase, succinyl-CoA synthase, LytB, DNA gyrase A, LepA, and Ser-tRNA synthetase) serve to distinguish the Rhizobiaceae, Brucellaceae, and Phyllobacteriaceae families from Bradyrhizobiaceae and Methylobacteriaceae. Based on the distribution patterns of these signatures, it is now possible to logically deduce a model for the branching order among alpha-proteobacteria, which is as follows: Rickettsiales --> Rhodospirillales-Sphingomonadales --> Rhodobacterales-Caulobacterales --> Rhizobiales (Rhizobiaceaea-Brucellaceae-Phyllobacteriaceae, and Bradyrhizobiaceae). The deduced branching order is also consistent with the topologies in the 16 rRNA and other phylogenetic trees. Signature sequences in a number of other proteins provide evidence that alpha-proteobacteria is a late branching taxa within Bacteria, which branched after the delta,epsilon-subdivisions but prior to the beta,gamma-proteobacteria. The shared presence of many of these signatures in the mitochondrial (eukaryotic) homologs also provides evidence of the alpha-proteobacterial ancestry of mitochondria.

Alphaproteobacteria↗

Detection and characterization of horizontal transfers in prokaryotes using genomic signature.

Horizontal DNA transfer is an important factor of evolution and participates in biological diversity. Unfortunately, the location and length of horizontal transfers (HTs) are known for very few species. The usage of short oligonucleotides in a sequence (the so-called genomic signature) has been shown to be species-specific even in DNA fragments as short as 1 kb. The genomic signature is therefore proposed as a tool to detect HTs. Since DNA transfers originate from species with a signature different from those of the recipient species, the analysis of local variations of signature along recipient genome may allow for detecting exogenous DNA. The strategy consists in (i) scanning the genome with a sliding window, and calculating the corresponding local signature (ii) evaluating its deviation from the signature of the whole genome and (iii) looking for similar signatures in a database of genomic signatures. A total of 22 prokaryote genomes are analyzed in this way. It has been observed that atypical regions make up approximately 6% of each genome on the average. Most of the claimed HTs as well as new ones are detected. The origin of putative DNA transfers is looked for among approximately 12 000 species. Donor species are proposed and sometimes strongly suggested, considering similarity of signatures. Among the species studied, Bacillus subtilis, Haemophilus Influenzae and Escherichia coli are investigated by many authors and give the opportunity to perform a thorough comparison of most of the bioinformatics methods used to detect HTs.

Bacillus subtilis↗

How to measure information carried by a modulated vocal signature?

Acoustic signaling systems that permit individual recognition are described in an increasing number of species. Evolutionary logic predicts that the efficiency of these signatures is related to the possibilities for confusion. To test this "signature adaptation" hypothesis, one needs a standardized method to estimate and compare the efficiency of different signatures. Beecher [Am. Zool. 22, 477-490 (1989)] developed such a method by comparing scalar parameters extracted from the signals. However, vocal signatures frequently consist in the evolution of one parameter against one other, which are not comparable through Beecher's method. Here we present a method to estimate the efficiency of modulated signatures. A signature's efficiency is given by its information capacity (Hm), derived from Shannon's information theory. The measure of Hm is based on an analysis of variance and uses the Euclidian distances between the signature's contours in the population. To validate our method, simulated datasets of modulated contours were used. The predicted efficiency of those signatures, estimated from Hm, was strongly correlated to its actual efficiency given by two classification methods: a discriminant analysis and a classification by human observers. Being also untied to sample size, Hm therefore allows comparing objectively vocal, but also visual and olfactory signatures.

Algorithms↗

Multiple robust signatures for detecting lymph node metastasis in head and neck cancer.

Genome-wide mRNA expression measurements can identify molecular signatures of cancer and are anticipated to improve patient management. Such expression profiles are currently being critically evaluated based on an apparent instability in gene composition and the limited overlap between signatures from different studies. We have recently identified a primary tumor signature for detection of lymph node metastasis in head and neck squamous cell carcinomas. Before starting a large multicenter prospective validation, we have thoroughly evaluated the composition of this signature. A multiple training approach was used for validating the original set of predictive genes. Based on different combinations of training samples, multiple signatures were assessed for predictive accuracy and gene composition. The initial set of predictive genes is a subset of a larger group of 825 genes with predictive power. Many of the predictive genes are interchangeable because of a similar expression pattern across the tumor samples. The head and neck metastasis signature has a more stable gene composition than previous predictors. Exclusion of the strongest predictive genes could be compensated by raising the number of genes included in the signature. Multiple accurate predictive signatures can be designed using various subsets of predictive genes. The absence of genes with strong predictive power can be compensated by including more genes with lower predictive power. Lack of overlap between predictive signatures from different studies with the same goal may be explained by the fact that there are more predictive genes than required to design an accurate predictor.

Carcinoma, Squamous Cell↗

Signature authentication by forensic document examiners.

We report on the first controlled study comparing the abilities of forensic document examiners (FDEs) and laypersons in the area of signature examination. Laypersons and professional FDEs were given the same signature-authentication/simulation-detection task. They compared six known signatures generated by the same person with six unknown signatures. No a priori knowledge of the distribution of genuine and nongenuine signatures in the unknown signature set was available to test-takers. Three different monetary incentive schemes were implemented to motivate the laypersons. We provide two major findings: (i) the data provided by FDEs and by laypersons in our tests were significantly different (namely, the hypothesis that there is no difference between the assessments provided by FDEs and laypersons about genuineness and nongenuineness of signatures was rejected); and (ii) the error rates exhibited by the FDEs were much smaller than those of the laypersons. In addition, we found no statistically significant differences between the data sets obtained from laypersons who received different monetary incentives. The most pronounced differences in error rates appeared when nongenuine signatures were declared authentic (Type I error) and when authentic signatures were declared nongenuine (Type II error). Type I error was made by FDEs in 0.49% of the cases, but laypersons made it in 6.47% of the cases. Type II error was made by FDEs in 7.05% of the cases, but laypersons made it in 26.1% of the cases.

Expert Testimony↗

Amino acid signatures in the normal cat retina.

PURPOSE: To establish a nomogram of amino acid signatures in normal neurons, glia, and retinal pigment epithelium (RPE) of the cat retina, guided by the premise that micromolecular signatures reflect cellular identity and metabolic integrity. The long-range objective was to provide techniques to detect subtle aberrations in cellular metabolism engendered by model interventions such as focal retinal detachment. METHODS: High-performance immunochemical mapping, image registration, and quantitative pattern recognition were combined to analyze the amino acid contents of virtually all cell types in serial 200-nm sections of normal cat retina. RESULTS: The cellular cohorts of the cat retina formed 14 separable biochemical theme classes. The photoreceptor --> bipolar cell --> ganglion cell pathway was composed of six classes, each possessing a characteristic glutamate signature. Amacrine cells could be grouped into two glycine- and three gamma-aminobutyric acid (GABA)-dominated populations. Horizontal cells possessed a distinctive GABA-rich signature completely separate from that of amacrine cells. A stable taurine-glutamine signature defined Müller cells, and a broad-spectrum aspartate-glutamate-taurine-glutamine signature was present in the normal RPE. CONCLUSIONS: In this study, basic micromolecular signatures were established for cat retina, and multiple metabolic subtypes were identified for each neurochemical class. It was shown that virtually all neuronal space can be accounted for by cells bearing characteristic glutamate, GABA, or glycine signatures. The resultant signature matrix constitutes a nomogram for assessing cellular responses to experimental challenges in disease models.

Alanine↗

Proteomic Signatures Related to Physical Activity Are Associated with Risks of Future Disease.

PURPOSE: Physical activity (PA) can lower the risk of developing chronic diseases. However, few studies have examined the proteomic signatures linked to PA, and the role of these signatures in the connection between PA levels and future disease risk remains unclear. This study aimed to investigate whether proteomic signatures indicative of PA are associated with the risk of developing common chronic diseases and to explore their role as statistical links in the relationship between PA levels and disease development. METHODS: We used data from a subcohort of UK Biobank participants. PA intensity data were collected from accelerometers worn by each participant. Plasma proteomics results were obtained through Olink analysis. The risks of developing each primary chronic disease were evaluated for types of PA and their associated proteomic signatures, adjusting for age, sex, ethnicity, socioeconomic status, lifestyle factors, and key measurement time-lag covariates. RESULTS: Based on the UK Biobank, we identified significant differences among the proteomic signatures of accelerometer-measured light PA, moderate-to-vigorous PA, and total PA. The main enriched pathways of these proteomic signatures included cell adhesion, cell migration, and immune response. Higher levels of accelerometer-measured PA and their associated proteomic signatures correlated with a lower risk of developing cardiometabolic disorders, cancers, psychological or neurological disorders, and respiratory diseases. CONCLUSIONS: Our findings show that PA and PA-related proteomic signatures are statistically associated with lower risks of chronic diseases. Further analyses identified proteins that were correlated with both PA and disease risk. These results need to be confirmed through longitudinal studies involving diverse populations.

Humans↗

Automated discovery of structural signatures of protein fold and function.

There are constraints on a protein sequence/structure for it to adopt a particular fold. These constraints could be either a local signature involving particular sequences or arrangements of secondary structure or a global signature involving features along the entire chain. To search systematically for protein fold signatures, we have explored the use of Inductive Logic Programming (ILP). ILP is a machine learning technique which derives rules from observation and encoded principles. The derived rules are readily interpreted in terms of concepts used by experts. For 20 populated folds in SCOP, 59 rules were found automatically. The accuracy of these rules, which is defined as the number of true positive plus true negative over the total number of examples, is 74% (cross-validated value). Further analysis was carried out for 23 signatures covering 30% or more positive examples of a particular fold. The work showed that signatures of protein folds exist, about half of rules discovered automatically coincide with the level of fold in the SCOP classification. Other signatures correspond to homologous family and may be the consequence of a functional requirement. Examination of the rules shows that many correspond to established principles published in specific literature. However, in general, the list of signatures is not part of standard biological databases of protein patterns. We find that the length of the loops makes an important contribution to the signatures, suggesting that this is an important determinant of the identity of protein folds. With the expansion in the number of determined protein structures, stimulated by structural genomics initiatives, there will be an increased need for automated methods to extract principles of protein folding from coordinates.

Algorithms↗

Correlated sequence-signatures as markers of protein-protein interaction.

As protein-protein interaction is intrinsic to most cellular processes, the ability to predict which proteins in the cell interact can aid significantly in identifying the function of newly discovered proteins, and in understanding the molecular networks they participate in. Here we demonstrate that characteristic pairs of sequence-signatures can be learned from a database of experimentally determined interacting proteins, where one protein contains the one sequence-signature and its interacting partner contains the other sequence-signature. The sequence-signatures that recur in concert in various pairs of interacting proteins are termed correlated sequence-signatures, and it is proposed that they can be used for predicting putative pairs of interacting partners in the cell. We demonstrate the potential of this approach on a comprehensive database of experimentally determined pairs of interacting proteins in the yeast Saccharomyces cerevisiae. The proteins in this database have been characterized by their sequence-signatures, as defined by the InterPro classification. A statistical analysis performed on all possible combinations of sequence-signature pairs has identified those pairs that are over-represented in the database of yeast interacting proteins. It is demonstrated how the use of the correlated sequence-signatures as identifiers of interacting proteins can reduce significantly the search space, and enable directed experimental interaction screens.

Computational Biology↗