Finding protein-binding sites in DNA sequences: the next generation.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to T Werner.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The detection of transcription control elements in DNA sequences became both more important and more complicated by the completion of the first full genome sequencing projects. Rapid evaluation of potential regulatory elements in large amounts of sequence data requires specific methods preferably available as user-friendly computer programs. However, many more algorithms and methods have been published than programs are available, creating problems for scientists who try to select an appropriate method for their needs from the literature. The Internet provides a worldwide and relatively easy access to computer software if the user knows where to look. One of the major problems remaining is how to find the appropriate software. We have compiled a guide detailing where software is available and what is to be expected in terms of interface and data compatibility with other programs. We also show results obtained with each program for several examples. The summarized features of each program should allow scientists to select quickly the method of their choice and inform them where to download the software.
UNLABELLED: In a randomized prospective study in 90 patients with COAD and tracheo-bronchial instability 3 groups were formed. Group 1: Therapy as group 3+ Physiotherapy with VRP1 Desitin, Group 2: Therapy as group 3+ Physiotherapy with RC-Cornet, Group 3: CONTROL GROUP: daily 40 mg prednisolon i.v., 2 x theophylline i.v. in relation to serum levels and 3 x inhalation of beta 2+ parasympathicolytic with a compressor inhaler. Therapy group 1 and 2 received the same drug and inhalation therapy as the controls. Controls of lung function before and after physiotherapy and visual analog scales for dyspnoea, cough, sputum and acceptance of the physiotherapy were performed at days 1, 4 and 7. With RC-Cornet the residual volume decreases statistically significant in comparison to VRP1 Desitin. Hyperventilation is also statistically significant smaller in RC-Cornet compared to VRP1 Desitin. The subjective improvement of sputum, dyspnoea and acceptance of the method of physiotherapy was statistically significant better for RC-Cornet. Regarding cough the significance was just failed by p < 0.055. RC-Cornet is a comfortable, effective, small accepted tool for the long term physiotherapy of patients with COAD and tracheobronchial instability.
In order to measure inotropic influences of physiologically occurring substances and drugs we used a newly developed guinea pig papillary muscle (GPPM) bioassay. GPPM were suspended in air and surface coated with buffer (Krebs-Henseleit solution). The muscles were stimulated (pulsating direct current, 1.5 V; 0.5 Hz, 20 ms duration) which led to contraction. This method enables measurements of inotropic effects up to 5 days, contrary to previous studies (1 day), in which immersions of GPPM in buffer were performed. In order to investigate the comparability of the new method we measured the effect of metabolites (citric acid cycle), lactic acid, lactate, and extracellular pH on muscle contractility. The H(+)-dependent decrease of the contractile force of the GPPM can be compensated by an increased Ca(2+)-concentration. Further, the influence of catecholamines (isoproterenol) on the contractility was investigated. As a result, isoproterenol caused arrhythmias and extrasystoles as it was observed in clinical studies. Several pharmaceutical substances were tested to show the reproducibility and repeatability of the bioassay.
Transcriptional control regions are usually composed of a complex arrangement of individual transcriptional elements like protein binding sites. This modular structure allows generation of enormous functional diversity of regulatory regions with a limited set of individual elements. We implemented simple formal representations of these general features of regulatory regions into an algorithm capable of developing complex models reflecting both the element composition and the functional organization of individual elements. Our method (ModelGenerator) requires a training set of at least 10 sequences containing the regulatory regions to be modelled and a very simple initial model which may consist of just two characteristic transcription factor binding sites. We show the capability of our algorithm to expand the initial model solely by comparative sequence analysis leading to complex, biologically meaningful models. A second program (ModelInspector) is capable to scan new sequence data for matches to models defined by ModelGenerator. We show two models for retroviral transcriptional control regions to be highly specific. A search against GenBank using one of the models is shown to be free of false negatives and to produce less than 2 false positives/million nucleotides. Thus, our algorithms appear to be useful tools for the analysis of extremely long genomic sequences which are now becoming available as results of various genome sequencing projects.
The gene coding for riboflavin synthase of Escherichia coli has been cloned by marker rescue on a 6-kb fragment that has been sequenced. The riboflavin synthase gene is identical to the ribC locus and codes for a protein of 213 amino acids with a mass of 23.4 kDa. It was mapped to a position at 37.5 min on the physical map of the E. coli chromosome. The 3' end of the ribC gene is directly adjacent to the cfa gene, which codes for cyclopropane-fatty-acid synthase. This gene is followed by two open reading frames designated ydhC and ydhB, which are predicted to code for putative proteins with 403 amino acids and 310 amino acids, respectively. The gene ydhC is similar to genes coding for resistance against various antibiotics (cmlA, bcr) and probably codes for a transmembrane protein. The protein specified by ydhB shows sequence similarity to a large family of DNA-binding proteins and probably represents a helix-turn-helix protein. The ydhB gene is directly adjacent to the regulatory gene purR. A 288-bp segment of the cfa gene has earlier been mapped incorrectly to a position adjacent to greA at 67 min. The ribC gene was hyperexpressed in recombinant E. coli strains to a level of about 30% of cellular protein. The protein was purified to homogeneity by chromatography. The specific activity was 26000 nmol.mg-1.h-1. The protein sediments at a velocity of S20 = 3.8 S. Sedimentation-equilibrium centrifugation indicated a molecular mass of 70 kDa, consistent with a trimer structure. The primary structure of riboflavin synthase is characterized by internal sequence similarity (25 identical amino acids in the C-terminal and N-terminal parts suggesting two structurally similar folding domains.
Explore the source record for details and available documents.
In this paper, a new way to think about, and to construct, pairwise as well as multiple alignments of DNA and protein sequences is proposed. Rather than forcing alignments to either align single residues or to introduce gaps by defining an alignment as a path running right from the source up to the sink in the associated dot-matrix diagram, we propose to consider alignments as consistent equivalence relations defined on the set of all positions occurring in all sequences under consideration. We also propose constructing alignments from whole segments exhibiting highly significant overall similarity rather than by aligning individual residues. Consequently, we present an alignment algorithm that (i) is based on segment-to-segment comparison instead of the commonly used residue-to-residue comparison and which (ii) avoids the well-known difficulties concerning the choice of appropriate gap penalties: gaps are not treated explicity, but remain as those parts of the sequences that do not belong to any of the aligned segments. Finally, we discuss the application of our algorithm to two test examples and compare it with commonly used alignment methods. As a first example, we aligned a set of 11 DNA sequences coding for functional helix-loop-helix proteins. Though the sequences show only low overall similarity, our program correctly aligned all of the 11 functional sites, which was a unique result among the methods tested. As a by-product, the reading frames of the sequences were identified. Next, we aligned a set of ribonuclease H proteins and compared our results with alignments produced by other programs as reported by McClure et al. [McClure, M. A., Vasi, T. K. & Fitch, W. M. (1994) Mol. Biol. Evol. 11, 571-592]. Our program was one of the best scoring programs. However, in contrast to other methods, our protein alignments are independent of user-defined parameters.
RFB virus is an ecotropic C-type retrovirus isolated from CF-1 mice, in which it is associated with induction of osteomas. Sequence analysis of the RFB provirus revealed no evidence for presence of an oncogene or a recombined env gene. RFB virus is a member of the murine leukemia virus (MuLV) group (RFB MuLV), sharing 97% nucleotide identity with the endogenous ecotropic provirus of AKR mice (Akv). Like Akv, expression of RFB MuLV mRNAs is inducible by dexamethasone treatment, indicating that FRB MuLV also shares transcriptional control signals with Akv. We assessed the pathogenic potential of RFB MuLV in NMRI mice, which, in contrast to CF-1 mice, do not contain endogenous ecotropic retroviruses. RFB MuLV induced osteomas, osteopetrosis, and lymphomas in newborn NMRI mice. Another CF-1 mouse-derived leukemia virus, FBJ MuLV, the helper virus of the FBJ osteosarcoma virus stock, as well as Akv, also induced osteomas, osteopetrosis, and lymphomas in NMRI mice similar to RFB MuLV. These findings indicate that endogenous retroviruses carry a pathogenic potential in hematopoietic tissues and in the skeleton.
Retroviruses are expressed under the control of viral control regions designated long terminal repeats (LTRs), which contain all signals for transcriptional initiation as well as transcriptional termination. However, retroviral LTRs from different species within a common genus, such as Lentivirus, do not show significant overall sequence homology. We compiled a model of the functional organization of 20 Lentivirus LTRs which we show to recognize all known Lentivirus LTRs. To this end we combined our previously published methods for identification of transcription elements with secondary structure element analysis in a novel modular approach. We deduced descriptions for three new Lentivirus-specific sequence elements present in most of the Lentivirus LTRs but absent in LTRs of other retrovirus families (B, C, D-type, BLV-HTLV, Spuma). Four of the 10 elements defined in our study were primate-specific. We were able to deduce a phylogeny based on our model which agrees in general with the phylogeny derived from the polymerase genes of these viruses. Our model indicated that more than 100 LTRs from the databases are of Lentivirus origin and can be clearly separated from all other LTR types (B, C, D, BLV-HTLV, Spuma). This selectivity appears to be a unique feature of our modular approach.
The GTP cyclohydrolase I (GTP-CH) gene of the cellular slime mould Dictyostelium discoideum has been cloned and sequenced. The 855 bp cDNA of this gene contains the open reading frame (ORF) encoding 232 amino acids with a predicted molecular mass of approx. 26 kDa. Southern blot analysis indicated the presence of a single gene for GTP-CH in Dictyostelium. PCR amplification of the ORF from chromosomal DNA and sequencing showed the existence of a 101 bp intron in the GTP-CH gene of Dictyostelium discoideum. The amino acid sequence has 47% and 49% positional identity to those of the human and yeast enzymes respectively. Most of the sequence variation between species is located in the N-terminal part of the protein. The overall identity with the E. coli protein is markedly lower. The enzyme was expressed in E. coli and purified as a 68 kDa fusion protein with the maltose-binding protein of E. coli. GTP-CH of Dictyostelium is heat-stable and showed maximal activity at 60 degrees C. The Km value for GTP is 50 microM.
The cDNA sequence of the beta B2-cry was determined from hamster (Mesocricetus auratus) and compared to the corresponding genes of bovine, frog, chicken, human, mouse and rat. Multispecies comparison demonstrated high homology between the hamster, rat and mouse gene, but larger distances to man, bovine, chicken and frog. There is striking identity within a strech of 36 deduced amino acids (aa) between the Greek key motif 3 and part of motif 4. This 36-aa domain contains a putative phosphorylation site for protein kinase C and is highly conserved among all known basic beta B-Cry; however, it can neither be detected in the acidic beta A-nor in the gamma-Cry.
We have identified a genomic clone containing the 5' regulatory region of the gene GTP-CH encoding human GTP cyclohydrolase I. The transcription start point (tsp) was mapped by 5'-rapid amplification of cDNA ends (5'-RACE). The 2.6-kb region upstream from the tsp showed promoter activity when ligated upstream from a reporter gene. The truncation of approximately 2 kb of the promoter did not change expression activity, while a further removal of 243 bp halved the activity. The promoter contains CCAAT and TATA boxes. The GC-rich region close to the tsp, which contains several putative Sp1-responsive elements, is required for maximum promoter activity. Interferon-gamma treatment of B-cells transfected with reporter constructs had no influence on the expression activity.
The speed of acquisition of genomic sequence data exceeds the evaluation of function of the sequences by a vast margin. Most software available for the prediction of individual features does not assess the correlation of different motifs (level 1 methods). Here, we present a second-level software package called GenomeInspector (GI) for further analysis of results obtained with level 1 methods. Our approach does not require any a priori knowledge about motif organization and was designed as a modular package with a graphical user interface. Three examples for GI application are presented.
The objective of this study was to assess the inhibitory effect of potential negative regulatory elements on human immunodeficiency virus (HIV)-1 long terminal repeat (LTR) activity. This was carried out by pairwise comparisons of reporter gene activities of HIV-LTR-CAT constructs differing in the presence and absence of nef sequences in transient transfection assays. Parallel transfections were performed in two persistently HIV-infected cell lines and the uninfected parental cell lines. The negative regulatory element (NRE) of the LTR did not suppress HIV LTR activity in any of the cell lines examined. However, a non-LTR-derived fragment of the nef gene had a distinct suppressive effect on activity of the full-length LTR in chronically infected astrocytoma cells. A weaker negative effect of this nef partial sequence (nps) was detected in the other cell lines with constructs lacking the NRE. The nps was capable of suppressing LTR activity in trans in chronically infected astrocytoma cells in a concentration-dependent manner. These results stress a negative role of a non-LTR nef partial sequence in a cell-specific manner. In addition, our data indicate that nps functions in trans with promoters unrelated to HIV LTR such as SV40 early and CMV immediate-early promoters.
Explore the source record for details and available documents.
We present an algorithm to identify potential functional elements like protein binding sites in DNA sequences, solely from nucleotide sequence data. Prerequisites are a set of at least seven not closely related sequences with a common biological function which is correlated to one or more unknown sequence elements present in most but not necessarily all of the sequences. The algorithm is based on a search for n-tuples which occur at least in a minimum percentage of the sequences with no or one mismatch, which may be at any position of the tuple. In contrast to functional tuples, random tuples show no preferred pattern of mismatch locations within the tuple nor is the conservation extended beyond the tuple. Both features of functional tuples are used to eliminate random tuples. Selection is carried out by maximization of the information content first for the n-tuple, then for a region containing the tuple and finally for the complete binding site. Further matches are found in an additional selection step, using the ConsInd method previously described. The algorithm is capable of identifying and delimiting elements (e.g. protein binding sites) represented by single short cores (e.g. TATA box) in sets of unaligned sequences of about 500 nucleotides using no information other than the nucleotide sequences. Furthermore, we show its ability to identify multiple elements in a set of complete LTR sequences (more than 600 nucleotides per sequence).