Search PubMed⌕ Search

Biomedical subjects

M S Gelfand

Publications and source records attributed to M S Gelfand.

At least 19 recordsLinked to original sources

Pro-Frame: similarity-based gene recognition in eukaryotic DNA sequences with errors.

Performance of existing algorithms for similarity-based gene recognition in eukaryotes drops when the genomic DNA has been sequenced with errors. A modification of the spliced alignment algorithm allows for gene recognition in sequences with errors, in particular frameshifts. It tolerates up to 5% of sequencing errors without considerable drop of prediction reliability when a sufficiently close homologous protein is available (normalized evolutionary distance similarity score 50% or higher).

Algorithms↗

Prediction of transcription regulatory sites in Archaea by a comparative genomic approach.

Intragenomic and intergenomic comparisons of upstream nucleotide sequences of archaeal genes were performed with the goal of predicting transcription regulatory sites (operators) and identifying likely regulons. Learning sets for the detection of regulatory sites were constructed using the available experimental data on archaeal transcription regulation or by analogy with known bacterial regulons, and further analysis was performed using iterative profile searches. The information content of the candidate signals detected by this method is insufficient for reliable predictions to be made. Therefore, this approach has to be complemented by examination of evolutionary conservation in different archaeal genomes. This combined strategy resulted in the prediction of a conserved heat shock regulon in all euryarchaea, a nitrogen fixation regulon in the methanogens Methanococcus jannaschii and Methanobacterium thermoautotrophicum and an aromatic amino acid regulon in M.thermoautotrophicum. Unexpectedly, the heat shock regulatory site was detected not only for genes that encode known chaperone proteins but also for archaeal histone genes. This suggests a possible function for archaeal histones in stress-related changes in DNA condensation. In addition, comparative analysis of the genomes of three Pyrococcus species resulted in the prediction of their purine metabolism and transport regulon. The results demonstrate the feasibility of prediction of at least some transcription regulatory sites by comparing poorly characterized prokaryotic genomes, particularly when several closely related genome sequences are available.

Archaea↗

ASDB: database of alternatively spliced genes.

Version 2.1 of ASDB (Alternative Splicing Data Base) contains 1922 protein and 2486 DNA sequences. The protein entries from SWISS-PROT are joined into clusters corresponding to alternatively spliced variants of one gene. The DNA division consists of complete genes with alternative splicing mentioned or annotated in GenBank. The search engine allows one to search over SWISS-PROT and GenBank fields and then follow the links to all variants. The database can be assessed at the URL http://cbcg.nersc.gov/asdb

Alternative Splicing↗

Transcriptional regulation of transport and utilization systems for hexuronides, hexuronates and hexonates in gamma purple bacteria.

The comparative approach is a powerful tool for the analysis of gene regulation in bacterial genomes. It can be applied to the analysis of regulons that have been studied experimentally as well as that of regulons for which no known regulatory sites are available. It is assumed that the set of co-regulated genes and the regulatory signal itself are conserved in related genomes. Here, we use genomic comparisons to study the regulation of transport and utilization systems for sugar acids in gamma purple bacteria Escherichia coli, Salmonella typhi, Klebsiella pneumoniae, Yersinia pestis, Erwinia chrysanthemi, Haemophilus influenzae and Vibrio cholerae. The variability of the operon structure and the location of the operator sites for the main transcription factors are demonstrated. The common metabolic map is combined with known and predicted regulatory interactions. It includes all known and predicted members of the GntR, UxuR/ExuR, KdgR, UidR and IdnR regulons. Moreover, most members of these regulons seem to be under catabolite repression mediated by CRP. The candidate UxuR/ExuR signal is proposed, the KdgR consensus is extended, and new operators for all transcription factors are identified in all studied genomes. Two new members of the KdgR regulon, a hypothetical ATP-dependent transport system OgtABCD and YjgK protein with unknown function, are detected. The former is likely to be the transport system for the products of pectin degradation, oligogalacturonides.

Amino Acid Sequence↗

Computer analysis of transcription regulatory patterns in completely sequenced bacterial genomes.

Recognition of transcription regulation sites (operators) is a hard problem in computational molecular biology. In most cases, small sample size and low degree of sequence conservation preclude the construction of reliable recognition rules. We suggest an approach to this problem based on simultaneous analysis of several related genomes. It appears that as long as a gene coding for a transcription regulator is conserved in the compared bacterial genomes, the regulation of the respective group of genes (regulons) also tends to be maintained. Thus a gene can be confidently predicted to belong to a particular regulon in case not only itself, but also its orthologs in other genomes have candidate operators in the regulatory regions. This provides for a greater sensitivity of operator identification as even relatively weak signals are likely to be functionally relevant when conserved. We use this approach to analyze the purine (PurR), arginine (ArgR) and aromatic amino acid (TrpR and TyrR) regulons of Escherichia coli and Haemophilus influenzae. Candidate binding sites in regulatory regions of the respective H.influenzae genes are identified, a new family of purine transport proteins predicted to belong to the PurR regulon is described, and probable regulation of arginine transport by ArgR is demonstrated. Differences in the regulation of some orthologous genes in E.coli and H.influenzae, in particular the apparent lack of the autoregulation of the purine repressor gene in H.influenzae, are demonstrated.

Arginine↗

ASDB: database of alternatively spliced genes.

A database of alternatively spliced genes (ASDB) has been constructed based on (i) the results of the analysis of Swiss-Prot entries containing products of these genes and (ii) clustering procedure joining proteins that could arise by alternative splicing of the same gene. ASDB incorporates information about alternatively spliced genes, their products and expression patterns. It can be searched in order to find all products of alternative splicing produced in a particular tissue or a given organism, or all variants generated by a particular transcript. ASDB currently contains about 1700 protein sequences and can be accessed via the Internet at URL http://cbcg.nersc.gov/asdb

Alternative Splicing↗

Statistical analysis of the exon-intron structure of higher and lower eukaryote genes.

Statistics of the exon-intron structure and splicing sites of several diverse eukaryotes was studied. The yeast exon-intron structures have a number of unique features. A yeast gene usually have at most one intron. The branch site is strongly conserved, whereas the polypirimidine tract is short. Long yeast introns tend to have stronger acceptor sites. In other species the branch site is less conserved and often cannot be determined. In non-yeast samples there is an almost universal correlation between lengths of neighboring exons (all samples excluding protists) and correlation between lengths of neighboring introns (human, drosophila, protists). On the average first introns are longer, and anomalously long introns are usually first introns in a gene. There is a universal preference for exons and exon pairs with the (total) length divisible by 3. Introns positioned between codons are preferred, whereas those positioned between the first and second positions in codon are avoided. The choice of A or G at the third position of intron (the donor splice sites generally prefer purines at this position) is correlated with the overall GC-composition of the gene. In all samples dinucleotide AG is avoided in the region preceding the acceptor site.

Algorithms↗

Segmentation of yeast DNA using hidden Markov models.

MOTIVATION: Compositionally homogeneous segments of genomic DNA often correspond to meaningful biological units. Simple sliding window analysis is usually insufficient for compositional segmentation of natural sequences. Hidden Markov models (HMM) with a small number of states are a natural language for description of compositional properties of chromosome-size DNA sequences. RESULTS: The algorithms were applied to yeast Saccharomyces cerevisiae chromosomes (YC) I, III, IV, VI and IX. The optimal number of HMM states is found to be four. The optimal four-state HMMs for all chromosomes are very similar, as well as the reconstructed segmentations. In most cases the models with k + 1 states are obtained by 'splitting' one of the states in the model with k states, and the corresponding increase of the level of detail in segmentation. The high AT states usually correspond to intergenic regions. We also explore the model's likelihood landscape and analyze the dynamics of the optimization process, thus addressing the problem of reliability of the obtained optima and efficiency of the algorithms.

Algorithms↗

Frequent alternative splicing of human genes.

Alternative splicing can produce variant proteins and expression patterns as different as the products of different genes, yet the prevalence of alternative splicing has not been quantified. Here the spliced alignment algorithm was used to make a first inventory of exon-intron structures of known human genes using EST contigs from the TIGR Human Gene Index. The results on any one gene may be incomplete and will require verification, yet the overall trends are significant. Evidence of alternative splicing was shown in 35% of genes and the majority of splicing events occurred in 5' untranslated regions, suggesting wide occurrence of alternative regulation. Most of the alternative splices of coding regions generated additional protein domains rather than alternating domains.

5' Untranslated Regions↗

Performance-guarantee gene predictions via spliced alignment.

An important and still unsolved problem in gene prediction is designing an algorithm that not only predicts genes but estimates the quality of individual predictions as well. Since experimental biologists are interested mainly in the reliability of individual predictions (rather than in the average reliability of an algorithm) we attempted to develop a gene recognition algorithm that guarantees a certain quality of predictions. We demonstrate here that the similarity level with a related protein is a reliable quality estimator for the spliced alignment approach to gene recognition. We also study the average performance of the spliced alignment algorithm for different targets on a complete set of human genomic sequences with known relatives and demonstrate that the average performance of the method remains high even for very distant targets. Using plant, fungal, and prokaryotic target proteins for recognition of human genes leads to accurate predictions with 95, 93, and 91% correlation coefficient, respectively. For target proteins with similarity score above 60%, not only the average correlation coefficient is very high (97% and up) but also the quality of individual predictions is guaranteed to be at least 82%. It indicates that for this level of similarity the worst case performance of the spliced alignment algorithm is better than the average case performance of many statistical gene recognition methods.

Algorithms↗

Algorithms and software for support of gene identification experiments.

MOTIVATION: Gene annotation is the final goal of gene prediction algorithms. However, these algorithms frequently make mistakes and therefore the use of gene predictions for sequence annotation is hardly possible. As a result, biologists are forced to conduct time-consuming gene identification experiments by designing appropriate PCR primers to test cDNA libraries or applying RT-PCR, exon trapping/amplification, or other techniques. This process frequently amounts to 'guessing' PCR primers on top of unreliable gene predictions and frequently leads to wasting of experimental efforts. RESULTS: The present paper proposes a simple and reliable algorithm for experimental gene identification which bypasses the unreliable gene prediction step. Studies of the performance of the algorithm on a sample of human genes indicate that an experimental protocol based on the algorithm's predictions achieves an accurate gene identification with relatively few PCR primers. Predictions of PCR primers may be used for exon amplification in preliminary mutation analysis during an attempt to identify a gene responsible for a disease. We propose a simple approach to find a short region from a genomic sequence that with high probability overlaps with some exon of the gene. The algorithm is enhanced to find one or more segments that are probably contained in the translated region of the gene and can be used as PCR primers to select appropriate clones in cDNA libraries by selective amplification. The algorithm is further extended to locate a set of PCR primers that uniformly cover all translated regions and can be used for RT-PCR and further sequencing of (unknown) mRNA.

Algorithms↗

Meningococcal cellulitis and sialadenitis.

Neisseria meningitidis is a rare cause of cellulitis. No cases of meningococcal sialadenitis have previously been reported. We recently successfully treated a patient who had meningococcal cellulitis and sialadenitis. We review previously reported cases of cellulitis due to N meningitidis and speculate on the role of underlying disease in the pathogenesis of this infection.

Aged↗

Avoidance of palindromic words in bacterial and archaeal genomes: a close connection with restriction enzymes.

Short palindromic sequences (4, 5 and 6 bp palindromes) are avoided at a statistically significant level in the genomes of several bacteria, including the completely sequenced Haemophilus influenzae and Synechocystis sp. genomes and in the complete genome of the archaeon Methanococcus jannaschii. In contrast, there is only moderate avoidance of palindromes in the small genome of the bacterium Mycoplasma genitalium and no detectable avoidance in the genomes of chloroplasts and mitochondria. The sites for type II restriction-modification enzymes detected in the given species tend to be among the most avoided palindromes in a particular genome, indicating a direct connection between the avoidance of short oligonucleotide words and restriction-modification systems with the respective specificity. Palindromes corresponding to sites for restriction enzymes from other species are also avoided, albeit less significantly, suggesting that in the course of evolution bacterial DNA has been exposed to a wide spectrum of restriction enzymes, probably as the result of lateral transfer mediated by mobile genetic elements, such as plasmids and prophages. Palindromic words appear to accumulate in DNA once it becomes isolated from restriction-modification systems, as demonstrated by the case of organellar genomes. By combining these observations with protein sequence analysis, we show that the most avoided 4-palindrome and the most avoided 6-palindrome in the archaeon M.jannaschii are likely to be recognition sites for two novel restriction-modification systems.

Archaea↗

Combinatorial approaches to gene recognition.

Recognition of genes via exon assembly approaches leads naturally to the use of dynamic programming. We consider the general graph-theoretical formulation of the exon assembly problem and analyze in detail some specific variants: multicriterial optimization in the case of non-linear gene-scoring functions; context-dependent schemes for scoring exons and related procedures for exon filtering; and highly specific recognition of arbitrary gene segments, oligonucleotide probes and polymerase chain reaction (PCR) primers.

Base Sequence↗

Encephalopathy due to capillary leak syndrome.

Systemic capillary leak syndrome (SCLS) is characterized by intermittent attacks of leakage of intravascular fluids into the extravascular space. Hypovolemia, hemoconcentration, weakness, edema, and visceral congestion are resulting manifestations of SCLS. Most patients with SCLS have clear mentation during attacks, and encephalopathy is not a known manifestation of the syndrome. We report a patient with acute idiopathic capillary leak syndrome manifested in an acute encephalopathy. The possibility of SCLS should be considered in patients who have an encephalopathy and hemoconcentration.

Brain Diseases↗