Search PubMedSearch

Biomedical subjects

A Viari

Publications and source records attributed to A Viari.

14 recordsLinked to original sources

Analysis of a Bacillus subtilis genome fragment using a co-operative computer system prototype.

Analysis of the huge volume of data generated by large scale sequencing projects requires the construction of new, sophisticated computer systems. These systems should be able to manage the biological data as well as the results of their analysis. They should also help the user to choose the most appropriate methods, and to string them together in order to solve a global analysis task. In this paper we present the prototype of a software system providing an environment for the analysis of large-scale sequence data. As a first step toward this end, this environment has been put to the test within the Bacillus subtilis genome sequencing project. This system integrates both the descriptive knowledge of the entities involved (genes, regulatory signals and the like) and the methodological knowledge comprising an extensible set of analytical methods. A knowledge representation based on two existing object-oriented models is used to implement this integrated system. In addition, the present prototype provides a suitable user interface both for displaying simultaneously the results generated by several methods and for interacting with the objects. We present in this paper the analysis of a B. subtilis genome fragment, present in data libraries but not annotated. Annotation of the genes present in the fragment allowed us to combine the results of several methods used for predicting coding sequences, and to characterize it as comprising a cryptic phage, the skin element. Comparison between the annotation of the skin element and a standard region of the chromosome indicated that local features of the nucleotide sequence could discriminate between phage and non-phage DNA sequence.

Bacillus subtilis

Finding flexible patterns in a text: an application to three-dimensional molecular matching.

Finding certain regularities in a text is an important problem in many areas, e.g. in the analysis of biological molecules such as nucleic acids or proteins. In the latter case, the text may be sequences of amino acids or a linear coding of three-dimensional structures, and the regularities then correspond to lexical or structural motifs common to two, or more, proteins. We first recall an earlier algorithm that found these regularities in a flexible way. Then we introduce a generalized version of this algorithm designed for the particular case of protein three-dimensional structures, since these structures present a few peculiarities that make them computationally harder to process. Finally, we give some applications of our new algorithm on concrete examples.

Algorithms

Cooperative computer system for genome sequence analysis.

Analysis of the huge volumes of data generated by large scale sequencing projects clearly requires the construction of new sophisticated computer systems. These systems should be able to handle the biological data as well as the results of the analysis of this data. They should also help the user to choose the most appropriate method for a simple task and to string together the methods needed to solve a global analysis task. In this paper we present the prototype of a software system that provides an environment for the analysis of large-scale sequence data. In a first approach this environment has been put to the test within the B. subtilis sequencing project. This system integrates both a descriptive knowledge of the entities involved (genes, regulatory signals etc.) and the methodological knowledge concerning an extendable set of analytical methods (i.e. how to solve a sequence analysis problem through task decomposition and method selection). A knowledge representation based on two existing object-oriented models, named Shirka and SCARP, is used to implement this integrated system. In addition, the present prototype provides a suitable user interface for both displaying the results generated by several methods and interacting with the objects. We present in this paper an overview of the knowledge-based models used to build this integrated system, and a description of the way in which biological entities and sequence analysis tasks are represented. We give illustrations of the co-operation between user and system during the problem solving process. Such a system constitutes a computer workbench for molecular biologists studying the genetic programs of living organisms.

Bacillus subtilis

A distance-based block searching algorithm.

We present in this paper an algorithm for the multiple comparison of a set of protein sequences. Our approach is that of peptide matching and consists in looking for all the words that occur approximatively in at least q of the sequences in the set, where q is a parameter. Words are compared by using a reference object called a model, that is itself a word over the alphabet of the amino acids, and the comparison between a model and a word is based on w-length words instead of single symbols. This idea is similar to the one used in the Blast program in the case of pairwise comparisons. Two w-length words are considered to be related if an alignment without gaps of the two using a similarity matrix has a score greater than a certain threshold value t. In our case, we say that a k-length word u is an occurrence of a model m of the same length if every w-length subword of u is related to the corresponding subword of m in the sense given above. If a model m has occurrences in at least q of the sequences of the set, m is said to occur in the set. In percentage terms, the value of q may correspond to something as small as 5% of the sequences (search for recurrent words in a set of non homologous proteins) or as high as 70-100% (establishment of a list of all similar words as a first step in a multiple alignment program). The algorithm presented here is an efficient and exact way of looking for all the models, of a fixed length k or of the greatest possible length kmax, that occur in a set of sequences. It can work with any kind of scoring matrix and an extension of the algorithm allows for the introduction of gaps between a model and its occurrences.

Algorithms

Saponins from Steganotaenia araliacea.

Six saponins have been isolated and identified from the leaves of Steganotaenia araliacea. They were identified as 3-O-[beta-D-galactopyranosyl(1----2)-(beta-D-galactopyranosyl (1----3))-beta-D-glucuronopyranosyl]-21-O-tigloyl and -21-O-angeloyl-R1-barrigenol, 3-O-[beta-D-glucopyranosyl(1----2)-(beta-D-xylopyranosyl (1----3))-beta-D-glucuronopyranosyl]-21-O-tigloyl and -21-O-angeloyl-R1-barrigenol, 3-O-[beta-D-glucopyranosyl(1----2)-(beta-D-glucopyranosyl-(1----3))-(alp ha-L- rhamnopyranosyl(1----4))-beta-D-glucopyranosyl] steganogenin and 3-O-[(beta-D-galactopyranosyl(1----2)-beta-D-glucuronopyranosyl]-2 8-O- beta-D-glucopyranosyl olean-12-ene-28-oic acid. Steganogenin is a new 17,22-seco-oleanolic acid derivative. The structures of the saponins were established by analysis of their 1H and 13C NMR spectra with the help of 2D-experiments and by Californium Plasma Desorption Mass Spectrometry.

Carbohydrate Sequence

Saponins from stem bark of Petersianthus macrocarpus.

Two bioactive saponins were isolated from the stem bark of Petersianthus macrocarpus. Their structures were elucidated by chemical degradations and by a combination of 2D NMR techniques and by Californium plasma desorption mass spectrometry. They are 3-O-([beta-D-galactopyranosyl (1-->2)][beta-D-galactopyranosyl (1-->3)]- beta-D-glucuronopyranosyl)-21-O-[3-(3-tigloyloxynilic acid)-4-tigloyloxy- alpha-L-arabinopyranosyl] barringtogenol C and 3-O-([beta-D-galactopyranosyl (1-->2)][beta-D-galactopyranosyl (1-->3)]-beta-D-glucuronopyranosyl)-28-O-alpha-L-rhamnopyranosyl barringtogenol C-21-O-benzoate. The absolute configuration of nilic acid was determined by partial synthesis. 3,3'-Dimethoxy ellagic acid and 3,3'-dimethoxy-4-O-beta-D- glucopyranosyl ellagic acid were also isolated.

Carbohydrate Sequence

Formation of cyclobutane thymine dimers photosensitized by pyridopsoralens: quantitative and qualitative distribution within DNA.

As after irradiation with 254-nm UV light, exposure of thymidine and three isomeric pyridopsoralen derivatives to UVA radiation, in the dry state, leads to the formation of the six diastereomers of cyclobutadithymidine as the predominant reaction. This unexpected photosensitized reaction, which also gives rise to both 5R* and 5S* diastereomers of 5,6-dihydro-5-(alpha-thymidylyl)thymidine (or "spore" photoproduct), is selective since [2 + 2] dimerization of 2'-deoxycytidine was not detected under the same experimental conditions. The cis-syn isomer of cyclobutadithymine was also found to be produced within isolated DNA following UVA irradiation in aqueous solutions containing 7-methylpyrido[3,4-c]psoralen. Quantitatively, this photoproduct represents about one-fifth of the overall yield of the furan-side pyridopsoralen [2 + 2] photocycloadducts to thymine. DNA sequencing methodology was used to demonstrate that pyridopsoralen-photosensitized DNA is a substrate for T4 endonuclease V and Escherichia coli photoreactivating enzyme, two enzymes acting specifically on cyclobutane pyrimidine dimers. Furthermore, the dimerization reaction of thymine is sequence dependent, with a different specificity from that mediated by far-UV irradiation as inferred from gel sequencing experiments. Interestingly, adjacent thymine residues are excellent targets for 7-methylpyrido[3,4-c]psoralen-mediated formation of cyclobutadithymine in TTTTA and TTAAT sites, which are also the strongest sites for photoaddition. The formation of cyclobutane thymine dimers concomitant to that of thymine-furocoumarin photoadducts and their eventual implication in the photobiological effects of the pyridopsoralens are discussed.

Autoradiography

'Multifrequency' location and clustering of sequence patterns from proteins.

In previous work, we have shown that a set of characteristics, defined as (code frequency) pairs, can be derived from a protein family by the use of a signal-processing method. This method enables the location and extraction of sequence patterns by taking into account each (code frequency) pair individually. In the present paper, we propose to extend this method in order to detect and visualize patterns by taking into account several pairs simultaneously. Two 'multifrequency' methods are described. The first one is based on a rewriting of the sequences with new symbols which summarize the frequency information. The second method is based on a clustering of the patterns associated with each pair. Both methods lead to the definition of significant consensus sequences. Some results obtained with calcium-binding proteins and serine proteases are also discussed.

Amino Acid Sequence

Escherichia coli molecular genetic map (1500 kbp): update II.

The DNA sequence data for Escherichia coli deposited in the EMBL library (release 27), together with miscellaneous data obtained from several laboratories, have been localized on an updated and corrected version of the restriction map of the chromosome generated by Kohara et al. (1987) and modified by others. This second update adds a further 500 kbp, increasing the amount of the E. coli chromosome sequenced to about one third of the total: 1510 kbp of sequenced DNA is included in the present data base. The accuracy of the map is assessed, and allows us to propose a precise genetic map position for every sequenced gene. The location of rare-cutting sites such as AvrII, NotI and SfiI have also been included in the update in order to combine the data obtained from different sources into one single file. The distribution of palindromic sequences (to which most restriction sites belong) has been studied in coding sequences. There appears to be a significant counter-selection against several such sequences in E. coli coding sequences (but not in other organisms such as Saccharomyces cerevisiae), suggesting the existence of constraints on DNA structure in E. coli, perhaps indicative of a functional role for horizontal gene transfer, preserving coding sequences, in this type of bacteria.

Base Sequence

Identification of the two cis-syn [2 + 2] cycloadducts resulting from the photoreaction of 3-carbethoxypsoralen with 2'-deoxycytidine and 2'-deoxyuridine.

Near-ultraviolet photolysis of 2'-deoxycytidine (dCyd) and 3-carbethoxypsoralen (3-CPs) in the dry state was found to generate two main stable photoadducts which were separated by thin-layer and high-performance liquid chromatography. Fast atom bombardment and plasma desorption mass spectrometry analyses suggested that the bound molecule to 3-CPs is dCyd. These two compounds were found to produce the corresponding 2'-deoxyuridine (dUrd) derivatives through a deamination process when left in aqueous solutions with a lifetime close to 24 h at 20 degrees C. The chemical structure of the deaminated photoadducts was confirmed by photochemical synthesis using dUrd as the substrate. UV and fluorescent measurements indicated that the furan moiety of 3-CPs is involved in the photobinding reaction. The cyclobutane type structure of the modified dUrd derivatives was established on the basis of its photoreversibility and detailed 1H NMR analysis. The cis-syn stereoconfiguration of the two photocycloadducts was inferred from coupling constant considerations and on the basis of the complete assignment of the cyclobutyl protons, requiring the synthesis of deuterated nucleosides at pyrimidine carbon C(6). Further confirmation of the diastereoisomeric relationship between the two cis-syn dUrd <54' 65'> 3-CPs was provided by circular dichroism measurements.

Deoxycytidine

A scale-independent signal processing method for sequence analysis.

In this paper, we present methods to detect and localize patterns in biologically related protein sequences (family). The patterns common to the sequences of the family are detected by using Fourier analysis. No previous scales (codes) are needed, they are actually produced as a result of the analysis procedure, together with the frequencies of the Fourier decompositions. Characteristic features of the family are thus expressed as (code-frequency) pairs. Various tools are proposed in order to localize the patterns, to compare the codes, and to evaluate the proximity of an arbitrary sequence to the investigated family. The general strategy is illustrated on a family composed of calcium-binding proteins.

Amino Acid Sequence

Plasma desorption mass spectrometric study of UV-induced lesions within DNA model compounds.

Positive and negative plasma desorption (PD) mass spectra of the di-deoxyribonucleoside monophosphate d(TpT) and of its biologically relevant ultraviolet-induced intramolecular photodimers are examined and discussed. The photodimers which were analysed by PD mass spectrometry include the cis-syn and trans-syn cyclobutyl isomers d(T[p]T), the pyrimidine-pyrimidone photoadduct (6-4)d(TpT) and its Dewar valence isomer. Molecular ions, quasi-molecular ions and several fragment ions are observed in all cases. It is shown that, despite the absence of mass differences between these dinucleoside monophosphates, the fragmentation pattern differs significantly between the two main classes: cyclobutane dimers and (6-4) adducts. PD mass spectrometry can therefore be envisaged for characterizing their formation within short DNA fragments.

DNA

Characterization and sequencing of normal and modified oligonucleotides by 252Cf plasma desorption mass spectrometry.

Plasma desorption (PD) mass spectra of normal deoxyribo-oligonucleotides and of neutral methylphosphonate deoxyribo-oligonucleosides are examined and discussed. Molecular ions of oligonucleotides up to nonamer have been observed for neutral species. It is also shown that PD mass spectra can be used to monitor chemical modifications of oligonucleotides, such as the covalent binding of an organic fluorescent probe, along the synthesis process.

Californium

A new class of psoralen photoadducts to DNA components: isolation and characterization of 8-MOP adducts to the osidic moiety of 2'-deoxyadenosine.

The near-UV-induced photoreaction of the bifunctional 8-methoxypsoralen (8-MOP) with 2'-deoxyadenosine (dAdo) was investigated in the dry state. Four main monoadducts of 8-MOP to 2'-deoxyadenosine were separated by high performance liquid chromatography and subsequently characterized by soft ionization mass spectrometry (fast atom bombardment and plasma desorption mass spectrometries) and extensive 1H NMR analysis including nuclear Overhauser effect (NOE) measurements. These new types of furocoumarin-nucleic acid component which appear to be specific to 2'-deoxyadenosine were shown to result from recombination of the 3,4-dihydropyron-4-yl radical of 8-MOP with 2'-deoxyadenosyl radical either at the 1' or the 5' position.

Chromatography, High Pressure Liquid