Search PubMed⌕ Search

Biomedical subjects

I Korf

Publications and source records attributed to I Korf.

8 recordsLinked to original sources

Integrating genomic homology into gene structure prediction.

TWINSCAN is a new gene-structure prediction system that directly extends the probability model of GENSCAN, allowing it to exploit homology between two related genomes. Separate probability models are used for conservation in exons, introns, splice sites, and UTRs, reflecting the differences among their patterns of evolutionary conservation. TWINSCAN is specifically designed for the analysis of high-throughput genomic sequences containing an unknown number of genes. In experiments on high-throughput mouse sequences, using homologous sequences from the human genome, TWINSCAN shows notable improvement over GENSCAN in exon sensitivity and specificity and dramatic improvement in exact gene sensitivity and specificity. This improvement can be attributed entirely to modeling the patterns of evolutionary conservation in genomic sequence.

Algorithms↗

Comparative genomic sequence analysis of the human and mouse cystic fibrosis transmembrane conductance regulator genes.

The identification of the cystic fibrosis transmembrane conductance regulator gene (CFTR) in 1989 represents a landmark accomplishment in human genetics. Since that time, there have been numerous advances in elucidating the function of the encoded protein and the physiological basis of cystic fibrosis. However, numerous areas of cystic fibrosis biology require additional investigation, some of which would be facilitated by information about the long-range sequence context of the CFTR gene. For example, the latter might provide clues about the sequence elements responsible for the temporal and spatial regulation of CFTR expression. We thus sought to establish the sequence of the chromosomal segments encompassing the human CFTR and mouse Cftr genes, with the hope of identifying conserved regions of biologic interest by sequence comparison. Bacterial clone-based physical maps of the relevant human and mouse genomic regions were constructed, and minimally overlapping sets of clones were selected and sequenced, eventually yielding approximately 1.6 Mb and approximately 358 kb of contiguous human and mouse sequence, respectively. These efforts have produced the complete sequence of the approximately 189-kb and approximately 152-kb segments containing the human CFTR and mouse Cftr genes, respectively, as well as significant amounts of flanking DNA. Analyses of the resulting data provide insights about the organization of the CFTR/Cftr genes and potential sequence elements regulating their expression. Furthermore, the generated sequence reveals the precise architecture of genes residing near CFTR/Cftr, including one known gene (WNT2/Wnt2) and two previously unknown genes that immediately flank CFTR/Cftr.

Animals↗

MaskerAid: a performance enhancement to RepeatMasker.

UNLABELLED: Identifying and masking repetitive elements is usually the first step when analyzing vertebrate genomic sequence. Current repeat identification software is sensitive but slow, creating a costly bottleneck in large-scale analyses. We have developed MaskerAid, a software enhancement to RepeatMasker that increased the speed of masking more than 30-fold at the most sensitive setting. AVAILABILITY: On request from the authors (see http://sapiens.wustl.edu/MaskerAid). CONTACT: maskeraid@watson.wustl.edu

Animals↗

MPBLAST : improved BLAST performance with multiplexed queries.

UNLABELLED: We have developed a program, MPBLAST, that increases the throughput of batch BLASTN searches by multiplexing (concatenating) query sequences and thereby reducing the number of actual database searches performed. Throughput was observed to increase in reciprocal proportion to the component sequence length. For sequencing read-sized queries of 500 bp, an order of magnitude speed-up was seen. AVAILABILITY: Free (see http://blast.wustl.edu) CONTACT: [ikorf, gish]@watson.wustl.edu

Computational Biology↗

The syntenic relationship of the zebrafish and human genomes.

The zebrafish is an important vertebrate model for the mutational analysis of genes effecting developmental processes. Understanding the relationship between zebrafish genes and mutations with those of humans will require understanding the syntenic correspondence between the zebrafish and human genomes. High throughput gene and EST mapping projects in zebrafish are now facilitating this goal. Map positions for 523 zebrafish genes and ESTs with predicted human orthologs reveal extensive contiguous blocks of synteny between the zebrafish and human genomes. Eighty percent of genes and ESTs analyzed belong to conserved synteny groups (two or more genes linked in both zebrafish and human) and 56% of all genes analyzed fall in 118 homology segments (uninterrupted segments containing two or more contiguous genes or ESTs with conserved map order between the zebrafish and human genomes). This work now provides a syntenic relationship to the human genome for the majority of the zebrafish genome.

Animals↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

The Polycomb group in Caenorhabditis elegans and maternal control of germline development.

Four Caenorhabditis elegans genes, mes-2, mes-3, mes-4 and mes-6, are essential for normal proliferation and viability of the germline. Mutations in these genes cause a maternal-effect sterile (i.e. mes) or grandchildless phenotype. We report that the mes-6 gene is in an unusual operon, the second example of this type of operon in C. elegans, and encodes the nematode homolog of Extra sex combs, a WD-40 protein in the Polycomb group in Drosophila. mes-2 encodes another Polycomb group protein (see paper by Holdeman, R., Nehrt, S. and Strome, S. (1998). Development 125, 2457-2467). Consistent with the known role of Polycomb group proteins in regulating gene expression, MES-6 is a nuclear protein. It is enriched in the germline of larvae and adults and is present in all nuclei of early embryos. Molecular epistasis results predict that the MES proteins, like Polycomb group proteins in Drosophila, function as a complex to regulate gene expression. Database searches reveal that there are considerably fewer Polycomb group genes in C. elegans than in Drosophila or vertebrates, and our studies suggest that their primary function is in controlling gene expression in the germline and ensuring the survival and proliferation of that tissue.

Amino Acid Sequence↗