Search PubMed⌕ Search

Biomedical subjects

Alexander Kozik

Publications and source records attributed to Alexander Kozik.

6 recordsLinked to original sources

Using variable rate models to identify genes under selection in sequence pairs: their validity and limitations for EST sequences.

Using likelihood-based variable selection models, we determined if positive selection was acting on 523 EST sequence pairs from two lineages of sunflower and lettuce. Variable rate models are generally not used for comparisons of sequence pairs due to the limited information and the inaccuracy of estimates of specific substitution rates. However, previous studies have shown that the likelihood ratio test (LRT) is reliable for detecting positive selection, even with low numbers of sequences. These analyses identified 56 genes that show a signature of selection, of which 75% were not identified by simpler models that average selection across codons. Subsequent mapping studies in sunflower show four of five of the positively selected genes identified by these methods mapped to domestication QTLs. We discuss the validity and limitations of using variable rate models for comparisons of sequence pairs, as well as the limitations of using ESTs for identification of positively selected genes.

Base Pairing↗

Analyses of synteny between Arabidopsis thaliana and species in the Asteraceae reveal a complex network of small syntenic segments and major chromosomal rearrangements.

Comparative genomic studies among highly divergent species have been problematic because reduced gene similarities make orthologous gene pairs difficult to identify and because colinearity is expected to be low with greater time since divergence from the last common ancestor. Nevertheless, synteny between divergent taxa in several lineages has been detected over short chromosomal segments. We have examined the level of synteny between the model species Arabidopsis thaliana and species in the Compositae, one of the largest and most diverse plant families. While macrosyntenic patterns covering large segments of the chromosomes are not evident, significant levels of local synteny are detected at a fine scale covering segments of 1-Mb regions of A. thaliana and regions of <5 cM in lettuce and sunflower. These syntenic patches are often not colinear, however, and form a network of regions that have likely evolved by duplications followed by differential gene loss.

Arabidopsis↗

High-density haplotyping with microarray-based expression and single feature polymorphism markers in Arabidopsis.

Expression microarrays hybridized with RNA can simultaneously provide both phenotypic (gene expression) and genotypic (marker) data. We developed two types of genetic markers from Affymetrix GeneChip expression data to generate detailed haplotypes for 148 recombinant inbred lines (RILs) derived from Arabidopsis thaliana accessions Bayreuth and Shahdara. Gene expression markers (GEMs) are based on differences in transcript levels that exhibit bimodal distributions in segregating progeny, while single feature polymorphism (SFP) markers rely on differences in hybridization to individual oligonucleotide probes. Unlike SFPs, GEMs can be derived from any type of DNA-based expression microarray. Our method identifies SFPs independent of a gene's expression level. Alleles for each GEM and SFP marker were ascertained with GeneChip data from parental accessions as well as RILs; a novel algorithm for allele determination using RIL distributions capitalized on the high level of genetic replication per locus. GEMs and SFP markers provided robust markers in 187 and 968 genes, respectively, which allowed estimation of gene order consistent with that predicted from the Col-0 genomic sequence. Using microarrays on a population to simultaneously measure gene expression variation and obtain genotypic data for a linkage map will facilitate expression QTL analyses without the need for separate genotyping. We have demonstrated that gene expression measurements from microarrays can be leveraged to identify polymorphisms across the genome and can be efficiently developed into genetic markers that are verifiable in a large segregating RIL population. Both marker types also offer opportunities for massively parallel mapping in unsequenced and less studied species.

Arabidopsis↗

Functional analysis of the plant disease resistance gene Pto using DNA shuffling.

Pto is a serine/threonine kinase that mediates resistance in tomato to strains of Pseudomonas syringae pv. tomato expressing the (a)virulence proteins AvrPto or AvrPtoB. DNA shuffling was used as a combinatorial in vitro genetic approach to dissect the functional regions of Pto. The Pto gene was shuffled with four of its paralogs from a resistant haplotype to create a library of recombinant products that was screened for interaction with AvrPto in yeast. All interacting clones and a representative sample of noninteracting clones were sequenced, and their ability to signal downstream was tested by the elicitation of a hypersensitive response in an AvrPto-dependent or -independent manner in planta. Eight candidate regions important for binding to AvrPto or for downstream signaling were identified by statistical correlations between individual amino acid positions and phenotype. A subset of the regions had previously been identified as important for recognition, confirming the validity of the shuffling approach. Three novel regions important for Pto function were validated by site-directed mutagenesis. Several chimeras and point mutants exhibited a differential interaction with (a)virulence proteins in the AvrPto and VirPphA family, demonstrating distinct binding requirements for different ligands. Additionally, the identification of chimeras that are both constitutively active as well as capable of binding AvrPto indicates that elicitation of downstream signaling does not involve a conformational change that precludes binding of AvrPto, as previously hypothesized. The correlations between phenotypes and variation generated by DNA shuffling paralleled natural variation observed between orthologs of Pto from Lycopersicon spp.

Blotting, Western↗

DiagHunter and GenoPix2D: programs for genomic comparisons, large-scale homology discovery and visualization.

The DiagHunter and GenoPix2D applications work together to enable genomic comparisons and exploration at both genome-wide and single-gene scales. DiagHunter identifies homologous regions (synteny blocks) within or between genomes. DiagHunter works efficiently with diverse, large datasets to predict extended and interrupted synteny blocks and to generate graphical and text output quickly. GenoPix2D allows interactive display of synteny blocks and other genomic features, as well as querying by annotation and by sequence similarity.

Algorithms↗

Genome-wide analysis of NBS-LRR-encoding genes in Arabidopsis.

The Arabidopsis genome contains approximately 200 genes that encode proteins with similarity to the nucleotide binding site and other domains characteristic of plant resistance proteins. Through a reiterative process of sequence analysis and reannotation, we identified 149 NBS-LRR-encoding genes in the Arabidopsis (ecotype Columbia) genomic sequence. Fifty-six of these genes were corrected from earlier annotations. At least 12 are predicted to be pseudogenes. As described previously, two distinct groups of sequences were identified: those that encoded an N-terminal domain with Toll/Interleukin-1 Receptor homology (TIR-NBS-LRR, or TNL), and those that encoded an N-terminal coiled-coil motif (CC-NBS-LRR, or CNL). The encoded proteins are distinct from the 58 predicted adapter proteins in the previously described TIR-X, TIR-NBS, and CC-NBS groups. Classification based on protein domains, intron positions, sequence conservation, and genome distribution defined four subgroups of CNL proteins, eight subgroups of TNL proteins, and a pair of divergent NL proteins that lack a defined N-terminal motif. CNL proteins generally were encoded in single exons, although two subclasses were identified that contained introns in unique positions. TNL proteins were encoded in modular exons, with conserved intron positions separating distinct protein domains. Conserved motifs were identified in the LRRs of both CNL and TNL proteins. In contrast to CNL proteins, TNL proteins contained large and variable C-terminal domains. The extant distribution and diversity of the NBS-LRR sequences has been generated by extensive duplication and ectopic rearrangements that involved segmental duplications as well as microscale events. The observed diversity of these NBS-LRR proteins indicates the variety of recognition molecules available in an individual genotype to detect diverse biotic challenges.

Amino Acid Sequence↗