Search PubMed⌕ Search

Biomedical subjects

Ian Korf

Publications and source records attributed to Ian Korf.

6 recordsLinked to original sources

Sex and tissue resolved co-expression networks reveal a female placental-brain axis protective against prenatal PCB exposure.

BACKGROUND: Neurodevelopmental disorders have a strong male bias that is poorly understood. The placenta provides molecular information about environmental interactions with genetics (including biological sex) that shape developmental processes in the brain. We investigate placental-brain transcriptional responses in an established mouse model of prenatal exposure to a human-relevant mixture of polychlorinated biphenyls (PCBs). RESULTS: To understand sex, tissue, and dosage effects in embryonic (E18) brain and placenta RNAseq data, we use weighted gene correlation network analysis (WGCNA) to create gene networks that could be compared across sex or tissue. WGCNA reveals that expression within most correlated gene networks is significantly and strongly associated with PCB exposure, but frequently in opposite directions between male-female and placenta-brain comparisons. In WGCNA and differentially expressed gene analyses, more transcriptional changes are observed in male brain than placenta, but the reverse is seen in females. Furthermore, female X-inactive specific transcript (Xist) levels correlate with sex-specific and non-monotonic PCB dose response, suggesting an X-linked protective epigenetic mechanism. The transcriptomic effects of low-dose PCB exposure are significantly opposed by dietary folic acid supplementation across both sexes but are strongest in female placentas. PCB and folic acid interacting gene networks are enriched in metabolic pathways involved in energy usage and translation, with female-specific protective effects enriched in PPAR, thermogenesis, glycerolipid, and O-glycan biosynthesis, as opposed to toxicant responses in male brain. CONCLUSIONS: A female protective effect in response to prenatal PCB exposure appears to be mediated by dose-dependent sex differences in transcriptional modulation of placental metabolic pathways.

Female↗

A probabilistic model of 3' end formation in Caenorhabditis elegans.

The 3' ends of mRNAs terminate with a poly(A) tail. This post-transcriptional modification is directed by sequence features present in the 3'-untranslated region (3'-UTR). We have undertaken a computational analysis of 3' end formation in Caenorhabditis elegans. By aligning cDNAs that diverge from genomic sequence at the poly(A) tract, we accurately identified a large set of true cleavage sites. When there are many transcripts aligned to a particular locus, local variation of the cleavage site over a span of a few bases is frequently observed. We find that in addition to the well-known AAUAAA motif there are several regions with distinct nucleotide compositional biases. We propose a generalized hidden Markov model that describes sequence features in C.elegans 3'-UTRs. We find that a computer program employing this model accurately predicts experimentally observed 3' ends even when there are multiple AAUAAA motifs and multiple cleavage sites. We have made available a complete set of polyadenylation site predictions for the C.elegans genome, including a subset of 6570 supported by aligned transcripts.

3' Untranslated Regions↗

Gene finding in novel genomes.

BACKGROUND: Computational gene prediction continues to be an important problem, especially for genomes with little experimental data. RESULTS: I introduce the SNAP gene finder which has been designed to be easily adaptable to a variety of genomes. In novel genomes without an appropriate gene finder, I demonstrate that employing a foreign gene finder can produce highly inaccurate results, and that the most compatible parameters may not come from the nearest phylogenetic neighbor. I find that foreign gene finders are more usefully employed to bootstrap parameter estimation and that the resulting parameters can be highly accurate. CONCLUSION: Since gene prediction is sensitive to species-specific parameters, every genome needs a dedicated gene finder.

Animals↗

Serial BLAST searching.

MOTIVATION: The translating BLAST algorithms are powerful tools for finding protein-coding genes because they identify amino acid similarities in nucleotide sequences. Unfortunately, these kinds of searches are computationally intensive and often represent bottlenecks in sequence analysis pipelines. Tuning parameters for speed can make the searches much faster, but one risks losing low-scoring alignments. However, high scoring alignments are relatively resistant to such changes in parameters, and this fact makes it possible to use a serial strategy where a fast, insensitive search is used to pre-screen a database for similar sequences, and a slow, sensitive search is used to produce the sequence alignments. RESULTS: Serial BLAST searches improve both the speed and sensitivity.

Algorithms↗

Leveraging the mouse genome for gene prediction in human: from whole-genome shotgun reads to a global synteny map.

The availability of draft sequences for both the mouse and human genomes makes it possible, for the first time, to annotate whole mammalian genomes using comparative methods. TWINSCAN is a gene-prediction system that combines the methods of single-genome predictors like GENSCAN with information derived from genome comparison, thereby improving accuracy. Because TWINSCAN uses genomic sequence only, it is less biased toward highly and/or ubiquitously expressed genes than GENEWISE, GENOMESCAN, and other methods based on evidence derived from transcripts. We show that TWINSCAN improves gene prediction in human using intermediate products from various stages of the sequencing and analysis of the mouse genome, from low-redundancy, whole-genome shotgun reads to the draft assembly and the synteny map. TWINSCAN improves on the prior state of the art even when alignments from only 1X coverage of the mouse genome are available. Gene prediction accuracy improves steadily from 1X through 3X, more slowly from 3X to 4X, and relatively little thereafter. The assembly and the synteny map greatly speed the computations, however. Our human annotation using the mouse assembly is conservative, predicting only 25,622 genes, and appears to be one of the best de novo annotations of the human genome to date.

Animals↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗