Search PubMed⌕ Search

Biomedical subjects

A Bairoch

Publications and source records attributed to A Bairoch.

At least 37 records · Page 2Linked to original sources

Improving protein identification from peptide mass fingerprinting through a parameterized multi-level scoring algorithm and an optimized peak detection.

We have developed a new algorithm to identify proteins by means of peptide mass fingerprinting. Starting from the matrix-assisted laser desorption/ionization-time-of-flight (MALDI-TOF) spectra and environmental data such as species, isoelectric point and molecular weight, as well as chemical modifications or number of missed cleavages of a protein, the program performs a fully automated identification of the protein. The first step is a peak detection algorithm, which allows precise and fast determination of peptide masses, even if the peaks are of low intensity or they overlap. In the second step the masses and environmental data are used by the identification algorithm to search in protein sequence databases (SWISS-PROT and/or TrEMBL) for protein entries that match the input data. Consequently, a list of candidate proteins is selected from the database, and a score calculation provides a ranking according to the quality of the match. To define the most discriminating scoring calculation we analyzed the respective role of each parameter in two directions. The first one is based on filtering and exploratory effects, while the second direction focuses on the levels where the parameters intervene in the identification process. Thus, according to our analysis, all input parameters contribute to the score, however with different weights. Since it is difficult to estimate the weights in advance, they have been computed with a generic algorithm, using a training set of 91 protein spectra with their environmental data. We tested the resulting scoring calculation on a test set of ten proteins and compared the identification results with those of other peptide mass fingerprinting programs.

Algorithms↗

A testis-specific gene, TPTE, encodes a putative transmembrane tyrosine phosphatase and maps to the pericentromeric region of human chromosomes 21 and 13, and to chromosomes 15, 22, and Y.

To contribute to the creation of a transcription map of human chromosome 21 (HC21) and to the identification of genes that may be involved in the pathogenesis of Down syndrome, exon trapping was performed from HC21-specific cosmids covering the entire chromosome. More than 700 exons have been identified to date. One such exon, hmc01a06, maps to YAC 831B6 which contains marker D21Z1 (alphoid repeats) and had previously been localized to the pericentromeric region of HC21. Northern-blot analysis revealed a 2.5-kb mRNA species strongly and exclusively expressed in the testis. We cloned the corresponding full-length cDNA, which encodes a predicted polypeptide of 551 amino acids with at least two potential transmembrane domains and a tyrosine phosphatase motif. The cDNA has sequence homology to chicken tensin, bovine auxilin and rat cyclin-G associated kinase (GAK). The entire polypeptide sequence also has significant homology to tumor suppressor PTEN/MMAC1 protein. We termed this novel gene/protein TPTE (transmembrane phosphatase with tensin homology). Polymerase chain reaction amplification, fluorescent in situ hybridization, Southern-blot and sequence analysis using monochromosomal somatic cell hybrids showed that this gene has highly homologous copies on HC13, 15, 22, and Y, in addition to its HC21 copy or copies. The estimated minimum number of copies of the TPTE gene in the haploid human genome is 7 in male and 6 in female. Zoo-blot analysis showed that TPTE is conserved between humans and other species. The biological function of the TPTE gene is presently unknown; however, its expression pattern, sequence homologies, and the presence of a potential tyrosine phosphatase domain suggest that it may be involved in signal transduction pathways of the endocrine or spermatogenetic function of the testis. It is also unknown whether all copies of TPTE are functional or whether some are pseudogenes. TPTE is, to our knowledge, the gene located closest to the human centromeric sequences.

Amino Acid Sequence↗

Eukaryotic aldehyde dehydrogenase (ALDH) genes: human polymorphisms, and recommended nomenclature based on divergent evolution and chromosomal mapping.

As currently being performed with an increasing number of superfamilies, a standardized gene nomenclature system is proposed here, based on divergent evolution, using multiple alignment analysis of all 86 eukaryotic aldehyde dehydrogenase (ALDH) amino-acid sequences known at this time. The ALDHs represent a superfamily of NAD(P)(+)-dependent enzymes having similar primary structures that oxidize a wide spectrum of endogenous and exogenous aliphatic and aromatic aldehydes. To date, a total of 54 animal, 15 plant, 14 yeast, and three fungal ALDH genes or cDNAs have been sequenced. These ALDHs can be divided into a total of 18 families (comprising 37 subfamilies), and all nonhuman ALDH genes are named here after the established human ALDH genes, when possible. An ALDH protein from one gene family is defined as having approximately < or = 40% amino-acid identity to that from another family. Two members of the same subfamily exhibit approximately > or = 60% amino-acid identity and are expected to be located at the same subchromosomal site. For naming each gene, it is proposed that the root symbol 'ALDH' denoting 'aldehyde dehydrogenase' be followed by an Arabic number representing the family and, when needed, a letter designating the subfamily and an Arabic number denoting the individual gene within the subfamily; all letters are capitalized in all mammals except mouse and fruit fly, e.g. 'human ALDH3A1 (mouse, Drosophila Aldh3a1).' It is suggested that the Human Gene Nomenclature Guidelines (http://++www.gene.ucl.ac.uk/nomenclature/guidelines.h tml) be used for all species other than mouse and Drosophila. Following these guidelines, the gene is italicized, whereas the corresponding cDNA, mRNA, protein or enzyme activity is written with upper-case letters and without italics, e.g. 'human, mouse or Drosophila ALDH3A1 cDNA, mRNA, or activity'. If an orthologous gene between species cannot be identified with certainty, sequential naming of these genes will be carried out in chronological order as they are reported to us. In addition, 20 human ALDH variant alleles that have been reported to date are listed herein and are recommended to be given numbers (or a number plus a capital letter) following an asterisk (e.g. 'ALDH3A2*2, ALDH2*4C'). It is anticipated that this eukaryotic ALDH gene nomenclature system will be extended to include bacterial genes within the next 2 years and that this nomenclature system will require updating on a regular basis; an ALDH Web site has been established for this purpose (http://++www.uchsc.edu/sp./sp./alcdbase/a ldhcov.html) and will serve as a medium for interaction amongst colleagues in this field.

Aldehyde Dehydrogenase↗

Protein identification with N and C-terminal sequence tags in proteome projects.

Genome sequences are available for increasing numbers of organisms. The proteomes (protein complement expressed by the genome) of many such organisms are being studied with two-dimensional (2D) gel electrophoresis. Here we have investigated the application of short N-terminal and C-terminal sequence tags to the identification of proteins separated on 2D gels. The theoretical N and C termini of 15, 519 proteins, representing all SWISS-PROT entries for the organisms Mycoplasma genitalium, Bacillus subtilis, Escherichia coli, Saccharomyces cerevisiae and human, were analysed. Sequence tags were found to be surprisingly specific, with N-terminal tags of four amino acid residues found to be unique for between 43% and 83% of proteins, and C-terminal tags of four amino acid residues unique for between 74% and 97% of proteins, depending on the species studied. Sequence tags of five amino acid residues were found to be even more specific. To utilise this specificity of sequence tags for protein identification, we created a world-wide web-accessible protein identification program, TagIdent (http://www.expasy.ch/www/tools.html), which matches sequence tags of up to six amino acid residues as well as estimated protein pI and mass against proteins in the SWISS-PROT database. We demonstrate the utility of this identification approach with sequence tags generated from 91 different E. coli proteins purified by 2D gel electrophoresis. Fifty-one proteins were unambiguously identified by virtue of their sequence tags and estimated pI and mass, and a further 11 proteins identified when sequence tags were combined with protein amino acid composition data. We conlcude that the TagIdent identification approach is best suited to the identification of proteins from prokaryotes whose complete genome sequences are available. The approach is less well suited to proteins from eukaryotes, as many eukaryotic proteins are not amenable to sequencing via Edman degradation, and tag protein identification cannot be unambiguous unless an organism's complete sequence is available.

Amino Acid Sequence↗

GPCRDB: an information system for G protein-coupled receptors.

The GPCRDB is a G protein-coupled receptor (GPCR) database system aimed at the collection and dissemination of GPCR related data. It holds sequences, mutant data and ligand binding constants as primary (experimental) data. Computationally derived data such as multiple sequence alignments, three dimensional models, phylogenetic trees and two dimensional visualization tools are added to enhance the database's usefulness. The GPCRDB is an EU sponsored project aimed at building a generic molecular class specific database capable of dealing with highly heterogeneous data. GPCRs were chosen as test molecules because of their enormous importance for medical sciences and due to the availability of so much highly heterogeneous data. The GPCRDB is available via the WWW at http://www.gpcr.org/7tm

Computer Communication Networks↗

Current status of the SWISS-2DPAGE database.

The SWISS-2DPAGE database (http: //www.expasy.ch/ch2d/ch2d-top.html ) consists of two-dimensional polyacrylamide gel electrophoresis images, as well as textual descriptions of the proteins that have been identified on them. The current release contains 15 reference maps from human biological samples, as well as from Saccharomyces cerevisiae , Escherichia coli and Dictyostelium discoideum origin. These reference maps have 2088 identified spots, corresponding to 410 separate protein entries in the database, in addition to virtual entries for each SWISS-PROT sequence.

Animals↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1998.

SWISS-PROT (http://www.expasy.ch/) is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Low molecular weight proteins: a challenge for post-genomic research.

The EcoGene project involves the examination of Escherichia coli K-12 DNA sequences and accompanying annotation in the public databases in order to refine the representation and prediction of the entire set of E. coli K-12 chromosomally encoded protein sequences. The results of this ongoing effort have been deposited in the SWISSPROT protein sequence database as sequencing of the E. coli genome has progressed to completion in recent years. Through this continuing research, we have discovered that the prediction of low molecular weight (small) proteins, arbitrarily defined as protein sequences < or = 150 amino acids (aa) in length, is problematic and requires special attention. We describe the small protein subset of EcoGene and the approach used to derive this subset from the complete E. coli genome sequence and database annotations. These E. coli proteins have helped to identify new small genes in other organisms and to identify conserved residues (motifs) using database searches and multiple alignments. Two thirds of the E. coli small proteins have not been characterized experimentally. The careful application of computer and laboratory methods to the analysis of small proteins is needed for accurate prediction, verification and characterization. The problem of accurate protein sequence identification is not limited to small proteins or to E. coli; these problems are encountered to varying degrees throughout all sequence databases.

Amino Acid Sequence↗

Two-dimensional gel electrophoresis for proteome projects: the effects of protein hydrophobicity and copy number.

Two-dimensional (2-D) gel electrophoresis is often used in proteome projects to provide a global view of the proteins expressed in any cell or tissue type. Here we have investigated the effects of protein hydrophobicity and cellular protein copy number on a protein's presence or absence on a two-dimensional gel. The average hydropathy values of all known proteins from Bacillus subtilis, Escherichia coli and Saccharomyces cerevisiae were calculated, thus defining the range of protein hydrophobicity and hydrophilicity in these organisms. The average hydropathy values were then calculated for a total of 427 proteins from these species, which had been identified elsewhere on 2-D gels. Strikingly, it was seen that no highly hydrophobic proteins, as defined by average hydrophobicity values, have been found to date on 2-D gel separations of whole cell lysates. A clear hydrophobicity cutoff point was seen, above which current 2-D electrophoresis methods appear not to be useful for protein separation. The effect of cellular protein copy number on a protein's presence on a 2-D gel was investigated by means of a graphical model. This model showed how variations in protein loading and copy number per cell interact to determine the quantity of a protein that will be present on a 2-D gel. Considering the current maximum in 2-D gel loading capacity, it was found that 2-D probably can not visualize or produce analytical quantities of proteins present at less than 1000 copies per cell. We conclude that further developments of 2-D electrophoresis techniques are desirable to enable the visualization and analysis of all proteins expressed by a cell or tissue.

Bacillus subtilis↗

'98 Escherichia coli SWISS-2DPAGE database update.

The combination of two-dimensional polyacrylamide gel electrophoresis (2-D PAGE), computer image analysis and several protein identification techniques allowed the Escherichia coli SWISS-2DPAGE database to be established. This is part of the ExPASy molecular biology server accessible through the WWW at the URL address http://www.expasy.ch/ch2d/ch2d-top.html . Here we report recent progress in the development of the E. coli SWISS-2DPAGE database. Proteins were separated with immobilized pH gradients in the first dimension and sodium dodecyl sulfate-polyacrylamide gel electrophoresis in the second dimension. To increase the resolution of the separation and thus the number of identified proteins, a variety of wide and narrow range immobilized pH gradients were used in the first dimension. Micropreparative gels were electroblotted onto polyvinylidene difluoride membranes and spots were visualized by amido black staining. Protein identification techniques such as amino acid composition analysis, gel comparison and microsequencing were used, as well as a recently described Edman "sequence tag" approach. Some of the above identification techniques were coupled with database searching tools. Currently 231 polypeptides are identified on the E. coli SWISS-2DPAGE map: 64 have been identified by N-terminal microsequencing, 39 by amino acid composition, and 82 by sequence tag. Of 153 proteins putatively identified by gel comparison, 65 have been confirmed. Many proteins have been identified using more than one technique. Faster progress in the E. coli proteome project will now be possible with advances in biochemical methodology and with the completion of the entire E. coli genome.

Bacterial Proteins↗

Multiple parameter cross-species protein identification using MultiIdent--a world-wide web accessible tool.

Recent increases in the number of genome sequencing projects means that the amount of protein sequence in databases is increasing at an astonishing pace. In proteome studies, this is facilitating the identification of proteins from molecularly well-defined organisms. However, in studies of proteins from the majority of organisms, proteins must be identified by comparing analytical data to sequences in databases from other species. This process is known as cross-species protein identification. Here we present a new program, MultiIdent, which uses multiple protein parameters such as amino acid composition, peptide masses, sequence tags, estimated protein pI and mass, to achieve cross-species protein identification. The program is structured so that protein amino acid composition, which is highly conserved across species boundaries, first generates a set of candidate proteins. These proteins are then queried with other protein parameters such as sequence tags and peptide masses. A final list of database entries which considers all analytical parameters is presented, ranked by an integrated score. We illustrate the power of the approach with the identification of a set of standard proteins, and the identification of proteins from dog heart separated by two-dimensional gel electrophoresis. The MultiIdent program is available on the world-wide web at: http://www.expasy.ch/sprot/multiident.h tml.

Amino Acid Sequence↗

A superfamily of metalloenzymes unifies phosphopentomutase and cofactor-independent phosphoglycerate mutase with alkaline phosphatases and sulfatases.

Sequence analysis of the probable archaeal phosphoglycerate mutase resulted in the identification of a superfamily of metalloenzymes with similar metal-binding sites and predicted conserved structural fold. This superfamily unites alkaline phosphatase, N-acetylgalactosamine-4-sulfatase, and cerebroside sulfatase, enzymes with known three-dimensional structures, with phosphopentomutase, 2,3-bisphosphoglycerate-independent phosphoglycerate mutase, phosphoglycerol transferase, phosphonate monoesterase, streptomycin-6-phosphate phosphatase, alkaline phosphodiesterase/nucleotide pyrophosphatase PC-1, and several closely related sulfatases. In addition to the metal-binding motifs, all these enzymes contain a set of conserved amino acid residues that are likely to be required for the enzymatic activity. Mutational changes in the vicinity of these residues in several sulfatases cause mucopolysaccharidosis (Hunter, Maroteaux-Lamy, Morquio, and Sanfilippo syndromes) and metachromatic leucodystrophy.

Alkaline Phosphatase↗

New insulin-like proteins with atypical disulfide bond pattern characterized in Caenorhabditis elegans by comparative sequence analysis and homology modeling.

We have identified three new families of insulin homologs in Caenorhabditis elegans. In two of these families, concerted mutations suggest that an additional disulfide bond links B and A domains, and that the A-domain internal disulfide bond is substituted by a hydrophobic interaction. Homology modeling remarkably confirms these predictions and shows that despite this atypical disulfide bond pattern and the absence of C-like peptide, all these proteins may adopt the same fold as the insulin. Interestingly, whereas we identified 10 insulin-like peptides, only one insulin-like-receptor (daf-2) has been found. We propose that these insulin-related peptides may correspond to different activators or inhibitors of the daf-2 insulin-regulating pathway.

Amino Acid Sequence↗