Search PubMed⌕ Search

Biomedical subjects

A Bairoch

Publications and source records attributed to A Bairoch.

At least 55 records · Page 3Linked to original sources

Construction of the gyrB Database for the Identification and Classification of Bacteria.

Nucleotide sequences of small-subunit rRNA (16S rRNA) are most commonly used for the identification and characterization of bacteria and their complex communities. However, 16S rRNA evolves slowly and is often not very convenient to resolve bacterial strains at the species level. We have therefore attempted to develop a rapid and more convenient system for bacterial identification using the gyrB gene sequences. We chose the gyrB gene, because (i) it is rarely transmitted horizontally, (ii) its molecular evolution rate is higher than that of 16S rRNA, and (iii) the gene is distributed ubiquitously among bacterial species. We PCR-amplified the 1.2 kb-long gyrB segments from about 1,000 bacterial species by using degenerate primers and determined their nucleotide sequences. The resultant data have been assembled into the gyrB database accessible via WWW.

Journal Article↗

Molecular basis of symbiosis between Rhizobium and legumes.

Access to mineral nitrogen often limits plant growth, and so symbiotic relationships have evolved between plants and a variety of nitrogen-fixing organisms. These associations are responsible for reducing 120 million tonnes of atmospheric nitrogen to ammonia each year. In agriculture, independence from nitrogenous fertilizers expands crop production and minimizes pollution of water tables, lakes and rivers. Here we present the complete nucleotide sequence and gene complement of the plasmid from Rhizobium sp. NGR234 that endows the bacterium with the ability to associate symbiotically with leguminous plants. In conjunction with transcriptional analyses, these data demonstrate the presence of new symbiotic loci and signalling mechanisms. The sequence and organization of genes involved in replication and conjugal transfer are similar to those of Agrobacterium, suggesting a recent lateral transfer of genetic information.

Bacterial Proteins↗

The PROSITE database, its status in 1997.

The PROSITE database consists of biologically significant patterns and profiles formulated in such a way that with appropriate computational tools it can help to determine to which known family of protein (if any) a new sequence belongs, or which known domain(s) it contains.

Amino Acid Sequence↗

The SWISS-PROT protein sequence data bank and its supplement TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, structure of its domains, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to two additional databases; a variety of new documentation files and the creation of TrEMBL, a computer annotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT.

Academies and Institutes↗

The UDP glycosyltransferase gene superfamily: recommended nomenclature update based on evolutionary divergence.

This review represents an update of the nomenclature system for the UDP glucuronosyltransferase gene superfamily, which is based on divergent evolution. Since the previous review in 1991, sequences of many related UDP glycosyltransferases from lower organisms have appeared in the database, which expand our database considerably. At latest count, in animals, yeast, plants and bacteria there are 110 distinct cDNAs/genes whose protein products all contain a characteristic 'signature sequence' and, thus, are regarded as members of the same superfamily. Comparison of a relatedness tree of proteins leads to the definition of 33 families. It should be emphasized that at least six cloned UDP-GlcNAc N-acetylglucosaminyltransferases are not sufficiently homologous to be included as members of this superfamily and may represent an example of convergent evolution. For naming each gene, it is recommended that the root symbol UGT for human (Ugt for mouse and Drosophila), denoting 'UDP glycosyltransferase,' be followed by an Arabic number representing the family, a letter designating the subfamily, and an Arabic numeral denoting the individual gene within the family or subfamily, e.g. 'human UGT2B4' and 'mouse Ugt2b5'. We recommend the name 'UDP glycosyltransferase' because many of the proteins do not preferentially use UDP glucuronic acid, or their nucleotide sugar preference is unknown. Whereas the gene is italicized, the corresponding cDNA, transcript, protein and enzyme activity should be written with upper-case letters and without italics, e.g. 'human or mouse UGT1A1.' The UGT1 gene (spanning > 500 kb) contains at least 12 promoters/first exons, which can be spliced and joined with common exons 2 through 5, leading to different N-terminal halves but identical C-terminal halves of the gene products; in this scheme each first exon is regarded as a distinct gene (e.g. UGT1A1, UGT1A2, ... UGT1A12). When an orthologous gene between species cannot be identified with certainty, as occurs in the UGT2B subfamily, sequential naming of the genes is being carried out chronologically as they become characterized. We suggest that the Human Gene Nomenclature Guidelines (http://www.gene.acl.ac.uk/nomenclature/guidelines.html++ +) be used for all species other than the mouse and Drosophila. Thirty published human UGT1A1 mutant alleles responsible for clinical hyperbilirubinemias are listed herein, and given numbers following an asterisk (e.g. UGT1A1*30) consistent with the Human Gene Nomenclature Guidelines. It is anticipated that this UGT gene nomenclature system will require updating on a regular basis.

Amino Acid Sequence↗

Arac/XylS family of transcriptional regulators.

The ArC/XylS family of prokaryotic positive transcriptional regulators includes more than 100 proteins and polypeptides derived from open reading frames translated from DNA sequences. Members of this family are widely distributed and have been found in the gamma subgroup of the proteobacteria, low- and high-G + C-content gram-positive bacteria, and cyanobacteria. These proteins are defined by a profile that can be accessed from PROSITE PS01124. Members of the family are about 300 amino acids long and have three main regulatory functions in common: carbon metabolism, stress response, and pathogenesis. Multiple alignments of the proteins of the family define a conserved stretch of 99 amino acids usually located at the C-terminal region of the regulator and connected to a nonconserved region via a linker. The conserved stretch contains all the elements required to bind DNA target sequences and to activate transcription from cognate promoters. Secondary analysis of the conserved region suggests that it contains two potential alpha-helix-turn-alpha-helix DNA binding motifs. The first, and better-fitting motif is supported by biochemical data, whereas existing biochemical data neither support nor refute the proposal that the second region possesses this structure. The phylogenetic relationship suggests that members of the family have recruited the nonconserved domain(s) into a series of existing domains involved in DNA recognition and transcription stimulation and that this recruited domain governs the role that the regulator carries out. For some regulators, it has been demonstrated that the nonconserved region contains the dimerization domain. For the regulators involved in carbon metabolism, the effector binding determinants are also in this region. Most regulators belonging to the AraC/XylS family recognize multiple binding sites in the regulated promoters. One of the motifs usually overlaps or is adjacent to the -35 region of the cognate promoters. Footprinting assays have suggested that these regulators protect a stretch of up to 20 bp in the target promoters, and multiple alignments of binding sites for a number of regulators have shown that the proteins recognize short motifs within the protected region.

Amino Acid Sequence↗

Protein sequence annotation in the genome era: the annotation concept of SWISS-PROT+TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation, a minimal level of redundancy and high level of integration with other databases. Ongoing genome sequencing projects have dramatically increased the number of protein sequences to be incorporated into SWISS-PROT. Since we do not want to dilute the quality standards of SWISS-PROT by incorporating sequences without proper sequence analysis and annotation, we cannot speed up the incorporation of new incoming data indefinitely. However, as we also want to make the sequences available as fast as possible, we introduced TREMBL (TRanslation of EMBL nucleotide sequence database), a supplement to SWISS-PROT. TREMBL consists of computer-annotated entries in SWISS-PROT format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except for CDS already included in SWISS-PROT. While TREMBL is already of immense value, its computer-generated annotation does not match the quality of SWISS-PROTs. The main difference is in the protein functional information attached to sequences. With this in mind, we are dedicating substantial effort to develop and apply computer methods to enhance the functional information attached to TREMBL entries.

Amino Acid Sequence↗

The PROSITE database, its status in 1995.

The PROSITE database consists of biologically significant patterns and profiles formulated in such a way that with appropriate computational tools it can help to determine to which known family of proteins (if any) a new sequence belongs or which known domain(s) it contains.

Computer Communication Networks↗

The SWISS-PROT protein sequence data bank and its new supplement TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc), a minimal level of redundancy and a high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to seven additional databases; a variety of new documentation files; the creation of TREMBL, and unannotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except CDS already included in SWISS-PROT.

Amino Acid Sequence↗

The ENZYME data bank in 1995.

The ENZYME data bank is a repository of information relative to the nomenclature of enzymes. The current version (October 1995) contains information relevant to 3594 enzymes. It is available from a variety of file and ftp servers as well as through the ExPASy World Wide Web server (http://expasy.hcuge.ch/).

CD-ROM↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. As in previous editions the genetic names are consistently associated to each sequence with a known and confirmed ORF. If necessary, synonyms are given in the case of allelic duplicated sequences. Although the first publication of a sequence gives-according to our rules-the genetic name of a gene, in some instances more commonly used names are given to avoid nomenclature problems and the use of ancient designations which are no longer used. In these cases the old designation is given as synonym. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, SWISSPROT and EMBL accession numbers. New entries will also contain the name from the systematic sequencing efforts. Since the release of LISTA4.1 we update the database continuously. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. This release includes reports from full Smith and Watermann peptide-level searches against a non-redundant protein sequence database. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS). The database is available by FTP and on World Wide Web.

Amino Acid Sequence↗

Federated two-dimensional electrophoresis database: a simple means of publishing two-dimensional electrophoresis data.

While a two-dimensional electrophoresis (2-DE) database is a relatively old concept, in recent years it generated renewed interest within the 2-DE community due to two main factors: (i) The high reproducibility of the current 2-DE method allows 2-DE images to be exchanged and compared between laboratories. (ii) The recent development of faster and more powerful techniques for protein identification such as microsequencing, matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS) and amino acid composition makes the production of reference protein maps and 2-DE databases cost- and time-effective. Additionally, the Internet network's current increase in popularity, combined with the rapid growth of Internet-connected laboratories, provides a straightforward means of publishing and sharing 2-DE data. While a small number of laboratories have already successfully published their data over the net, the increasing number of 2-DE database servers that are currently being set up will sooner or later require some kind of standardization. Unfortunately, standardization can be a long and cumbersome process inevitably leading to undesirable compromises. A federated database offers a simple and efficient way to publish and share 2-DE data without the need for standardization. Taking advantage of Internet protocols such as World Wide Web, they allow each laboratory to maintain their own database and to interconnect it with other similar databases through the use of active cross-references. This paper first presents guidelines for building a federated 2-DE database that may easily be followed by most laboratories. It then briefly reviews the state-of-the-art in networked 2-DE databases, and finally describes the SWISS-2DPAGE database which fully implements the concept of a federated 2-DE database.

Computer Communication Networks↗

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence↗