Search PubMed⌕ Search

Biomedical subjects

V Solovyev

Publications and source records attributed to V Solovyev.

5 recordsLinked to original sources

An integrated gene annotation and transcriptional profiling approach towards the full gene content of the Drosophila genome.

BACKGROUND: While the genome sequences for a variety of organisms are now available, the precise number of the genes encoded is still a matter of debate. For the human genome several stringent annotation approaches have resulted in the same number of potential genes, but a careful comparison revealed only limited overlap. This indicates that only the combination of different computational prediction methods and experimental evaluation of such in silico data will provide more complete genome annotations. In order to get a more complete gene content of the Drosophila melanogaster genome, we based our new D. melanogaster whole-transcriptome microarray, the Heidelberg FlyArray, on the combination of the Berkeley Drosophila Genome Project (BDGP) annotation and a novel ab initio gene prediction of lower stringency using the Fgenesh software. RESULTS: Here we provide evidence for the transcription of approximately 2,600 additional genes predicted by Fgenesh. Validation of the developmental profiling data by RT-PCR and in situ hybridization indicates a lower limit of 2,000 novel annotations, thus substantially raising the number of genes that make a fly. CONCLUSIONS: The successful design and application of this novel Drosophila microarray on the basis of our integrated in silico/wet biology approach confirms our expectation that in silico approaches alone will always tend to be incomplete. The identification of at least 2,000 novel genes highlights the importance of gathering experimental evidence to discover all genes within a genome. Moreover, as such an approach is independent of homology criteria, it will allow the discovery of novel genes unrelated to known protein families or those that have not been strictly conserved between species.

Animals↗

A novel type of RNase III family proteins in eukaryotes.

The RNase III family of double-stranded RNA-specific endonucleases is characterized by the presence of a highly conserved 9 amino acid stretch in their catalytic center known as the RNase III signature motif. We isolated the drosha gene, a new member of this family in Drosophila melanogaster. Characterization of this gene revealed the presence of two RNase III signature motifs in its sequence that may indicate that it is capable of forming an active catalytic center as a monomer. The drosha protein also contains an 825 amino acid N-terminus with an unknown function. A search for the known homologues of the drosha protein revealed that it has a similarity to two adjacent annotated genes identified during C. elegans genome sequencing. Analysis of the genomic region of these genes by the Fgenesh program and sequencing of the EST cDNA clone derived from it revealed that this region encodes only one gene. This newly identified gene in nematode genome shares a high similarity to Drosophila drosha throughout its entire protein sequence. A potential drosha homologue is also found among the deposited human cDNA sequences. A comparison of these drosha proteins to other members of the RNase III family indicates that they form a new group of proteins within this family.

Amino Acid Sequence↗

The Gene-Finder computer tools for analysis of human and model organisms genome sequences.

We present a complex of new programs for promoter, 3'-processing, splice sites, coding exons and gene structure identification in genomic DNA of several model species. The human gene structure prediction program FGENEH, exon prediction-FEXH and splice site prediction-HSPL have been modified for sequence analysis of Drosophila (FGENED, FEXD and DSPL), C.elegance (FGENEN, FEXN and NSPL), Yeast (FEXY and YSPL) and Plant (FGENEA, FEXA and ASPL) genomic sequences. We recomputed all frequency and discriminant function parameters for these organisms and adjusted organism specific minimal intron lengths. An accuracy of coding region prediction for these programs is similar with the observed accuracy of FEXH and FGENEH. We have developed FEXHB and FGENEHB programs combining pattern recognition features and information about similarity of predicted exons with known sequences in protein databases. These programs have approximately 10% higher average accuracy of coding region recognition. Two new programs for human promoter site prediction (TSSG and TSSW) have been developed which use Gosh (1993) and Wingender (1994) data bases of functional motifs, respectively. POLYAH program was designed for prediction of 3'-processing regions in human genes and CDSB program was developed for bacterial gene prediction. We have developed a new approach to predict multiple genes based on double dynamic programming, that is very important for analysis of long genomic DNA fragments generated by genome sequencing projects. Analysis of uncharacterized sequences based on our methods is available through the University of Houston, Weizmann Institute of Science email servers and several Web pages at Baylor College of Medicine.

Animals↗

Expression of msl-2 causes assembly of dosage compensation regulators on the X chromosomes and female lethality in Drosophila.

Male-specific lethal-2 (msl-2) is a RING finger protein that is required for X chromosome dosage compensation in Drosophila males. Consistent with the formation of a dosage compensation protein complex, msl-2 colocalizes with the other MSL proteins on the male X chromosome and coimmunoprecipitates with msl-1 from male larval extracts. Ectopic expression of msl-2 in females results in the appearance of the other MSL dosage compensation regulators on the female X chromosomes and decreased female viability. We suggest that msl-2 RNA is the primary target of SxI regulation in the dosage compensation pathway and present a speculative model for the regulation of two distinct modes of dosage compensation by SxI.

Amino Acid Sequence↗

Integrated databases and computer systems for studying eukaryotic gene expression.

MOTIVATION: The goal of the work was to develop a WWW-oriented computer system providing a maximal integration of informational and software resources on the regulation of gene expression and navigation through them. Rapid growth of the variety and volume of information accumulated in the databases on regulation of gene expression necessarily requires the development of computer systems for automated discovery of the knowledge that can be further used for analysis of regulatory genomic sequences. RESULTS: The GeneExpress system developed includes the following major informational and software modules: (1) Transcription Regulation (TRRD) module, which contains the databases on transcription regulatory regions of eukaryotic genes and TRRD Viewer for data visualization; (2) Site Activity Prediction (ACTIVITY), the module for analysis of functional site activity and its prediction; (3) Site Recognition module, which comprises (a) B-DNA-VIDEO system for detecting the conformational and physicochemical properties of DNA sites significant for their recognition, (b) Consensus and Weight Matrices (ConsFrec) and (c) Transcription Factor Binding Sites Recognition (TFBSR) systems for detecting conservative contextual regions of functional sites and their recognition; (4) Gene Networks (GeneNet), which contains an object-oriented database accumulating the data on gene networks and signal transduction pathways, and the Java-based Viewer for exploration and visualization of the GeneNet information; (5) mRNA Translation (Leader mRNA), designed to analyze structural and contextual properties of mRNA 5'-untranslated regions (5'-UTRs) and predict their translation efficiency; (6) other program modules designed to study the structure-function organization of regulatory genomic sequences and regulatory proteins. AVAILABILITY: GeneExpress is available at http://wwwmgs.bionet.nsc. ru/systems/GeneExpress/ and the links to the mirror site(s) can be found at http://wwwmgs.bionet.nsc.ru/mgs/links/mirrors.html+ ++.

Algorithms↗