Search PubMed⌕ Search

Biomedical subjects

Anton Nekrutenko

Publications and source records attributed to Anton Nekrutenko.

16 recordsLinked to original sources

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software↗

Functionality of unspliced XBP1 is required to explain evolution of overlapping reading frames.

Eukaryotic genes with overlapping reading frames exemplify some of the most striking biological phenomena. Transcript of one such gene, the gene encoding X-box protein (XBP1), has evolved a mechanism for "on-demand" switching of translation between two overlapping reading frames. Despite the existence of this elaborate system, only one reading frame was believed to be functional. Here, we show that XBP1 evolves in a fashion that is only consistent with functionality of both reading frames. Our study provides a novel evolutionary framework for the analysis of loci with overlapping reading frames, which can be used for identification and analyses of novel dual-coding genes.

Animals↗

mNSC1 shows no evidence of protein-coding capacity.

The Mus musculus non-selective cation channel gene mNSC1 was used as a classical example of a gene derived from transposable elements. To study the evolution of mNSC1 in M. musculus we sequenced this locus in M. musculus, M. hortulanus, M. spretus, M. caroli, and M. pahari. We found that the previously published 1,275 bp coding region was not present in any of these species. We identified a second possible coding region that was present only in M. musculus. However, RT-PCR experiments did not confirm the expression of the second reading frame. Our findings suggest that mNSC1 lacks protein-coding capacity and highlight the need for comparative validation of all TE-containing genes.

Animals↗

Rapid and asymmetric divergence of duplicate genes in the human gene coexpression network.

BACKGROUND: While gene duplication is known to be one of the most common mechanisms of genome evolution, the fates of genes after duplication are still being debated. In particular, it is presently unknown whether most duplicate genes preserve (or subdivide) the functions of the parental gene or acquire new functions. One aspect of gene function, that is the expression profile in gene coexpression network, has been largely unexplored for duplicate genes. RESULTS: Here we build a human gene coexpression network using human tissue-specific microarray data and investigate the divergence of duplicate genes in it. The topology of this network is scale-free. Interestingly, our analysis indicates that duplicate genes rapidly lose shared coexpressed partners: after approximately 50 million years since duplication, the two duplicate genes in a pair have only slightly higher number of shared partners as compared with two random singletons. We also show that duplicate gene pairs quickly acquire new coexpressed partners: the average number of partners for a duplicate gene pair is significantly greater than that for a singleton (the latter number can be used as a proxy of the number of partners for a parental singleton gene before duplication). The divergence in gene expression between two duplicates in a pair occurs asymmetrically: one gene usually has more partners than the other one. The network is resilient to both random and degree-based in silico removal of either singletons or duplicate genes. In contrast, the network is especially vulnerable to the removal of highly connected genes when duplicate genes and singletons are considered together. CONCLUSION: Duplicate genes rapidly diverge in their expression profiles in the network and play similar role in maintaining the network robustness as compared with singletons.

Chromosome Mapping↗

Galaxy: a platform for interactive large-scale genome analysis.

Accessing and analyzing the exponentially expanding genomic sequence and functional data pose a challenge for biomedical researchers. Here we describe an interactive system, Galaxy, that combines the power of existing genome annotation databases with a simple Web portal to enable users to search remote resources, combine data from independent queries, and visualize the results. The heart of Galaxy is a flexible history system that stores the queries from each user; performs operations such as intersections, unions, and subtractions; and links to other computational tools. Galaxy can be accessed at http://g2.bx.psu.edu.

Biological Evolution↗

Oscillating evolution of a mammalian locus with overlapping reading frames: an XLalphas/ALEX relay.

XLalphas and ALEX are structurally unrelated mammalian proteins translated from alternative overlapping reading frames of a single transcript. Not only are they encoded by the same locus, but a specific XLalphas/ALEX interaction is essential for G-protein signaling in neuroendocrine cells. A disruption of this interaction leads to abnormal human phenotypes, including mental retardation and growth deficiency. The region of overlap between the two reading frames evolves at a remarkable speed: the divergence between human and mouse ALEX polypeptides makes them virtually unalignable. To trace the evolution of this puzzling locus, we sequenced it in apes, Old World monkeys, and a New World monkey. We show that the overlap between the two reading frames and the physical interaction between the two proteins force the locus to evolve in an unprecedented way. Namely, to maintain two overlapping protein-coding regions the locus is forced to have high GC content, which significantly elevates its intrinsic evolutionary rate. However, the two encoded proteins cannot afford to change too quickly relative to each other as this may impair their interaction and lead to severe physiological consequences. As a result XLalphas and ALEX evolve in an oscillating fashion constantly balancing the rates of amino acid replacements. This is the first example of a rapidly evolving locus encoding interacting proteins via overlapping reading frames, with a possible link to the origin of species-specific neurological differences.

Amino Acid Sequence↗

Reconciling the numbers: ESTs versus protein-coding genes.

The number of expressed sequences greatly surpasses the estimated number of protein-coding genes in mammalian genomes. An evolutionary approach reveals that only 9% to 14% of human-expressed and mouse-expressed sequences are able to code for proteins. Clustering of these sequences using cross-species relationships suggests that millions of expressed sequences may correspond to only approximately 20,000 distinct protein-coding transcripts.

Animals↗

Identification of novel exons from rat-mouse comparisons.

Exon shuffling, a major mechanism of gene evolution, scrambles existing sequences to create new genes. However, is it possible for an exon to be created from scratch? Here we conduct a survey of rat and mouse genomes and identify 2302 putative rodent-specific exons absent from the human genome. Analysis of rodent transcripts supporting these exons indicates that over half appear to be alternatively spliced in genes orthologous between rodents and human. This study demonstrates the importance of sequencing genomes from multiple species to accurately document the evolution of gene structure.

Animals↗

Comparative genomics.

The genomes from three mammals (human, mouse, and rat), two worms, and several yeasts have been sequenced, and more genomes will be completed in the near future for comparison with those of the major model organisms. Scientists have used various methods to align and compare the sequenced genomes to address critical issues in genome function and evolution. This review covers some of the major new insights about gene content, gene regulation, and the fraction of mammalian genomes that are under purifying selection and presumed functional. We review the evolutionary processes that shape genomes, with particular attention to variation in rates within genomes and along different lineages. Internet resources for accessing and analyzing the treasure trove of sequence alignments and annotations are reviewed, and we discuss critical problems to address in new bioinformatic developments in comparative genomics.

Animals↗

ETOPE: Evolutionary test of predicted exons.

Since a large number of computationally predicted exons are not supported by existing sequence (e.g. ESTs) or experimental (e.g. expression analysis) data they need to be validated by other methods. ETOPE is designed to test computational predictions by using signals that have not been included in any current computational prediction method. The test is based on the ratio of non-synonymous to synonymous substitution rates between sequences from different genomes. It has been previously shown, by empirical data and computer simulation, to be a powerful criterion for identifying protein-coding regions. The ETOPE is available at http://nekrut.uchicago.edu/etope/.

Amino Acid Substitution↗

Evolutionary dynamics of oncogenes and tumor suppressor genes: higher intensities of purifying selection than other genes.

Oncogenes and tumor suppressor genes (hereafter referred to as "cancer genes") result in cancer when they experience substitutions that prevent or distort their normal function. We examined evolutionary pressures acting on cancer genes and other classes of disease-related genes and compared our results to analyses of genes without known association to disease. We compared synonymous and nonsynonymous substitution rates in 3,035 human genes-approximately 10% of the genome-measuring the intensity of purifying selection on 311 human disease genes, including 122 cancer-related genes. Although the genes examined are similar to nondisease genes in product, expression, function, and pathway affiliation, we found intriguing differences in the selective pressures experienced by cancer genes relative to other (noncancer) disease-related and non-disease-related genes. We found a statistically significant increase in the intensity of purifying selection exerted on cancer genes (the average ratio of nonsynonymous to synonymous substitutions, omega, was 0.079) relative to all other disease-related genes groups (omega = 0.101) and non-disease-related genes (omega = 0.100). This difference indicates a striking increase in selection against nonsynonymous substitutions in oncogenes and tumor suppressor genes. This finding provides insight into the etiology of cancer and the differences between genes involved in cancer and those implicated in other human diseases. Specifically, we found a significant overlap between human oncogenes and tumor suppressor genes and "essential genes," human homologs of mouse lethal genes identified by knockout experiments. This insight may improve our ability to identify cancer-related genes and enhances our understanding of the nature of these genes.

Evolution, Molecular↗

Subgenome-specific markers in allopolyploid cotton Gossypium hirsutum: implications for evolutionary analysis of polyploids.

We developed a set of genetic markers specific to the A and D genome types of cotton using representational difference analysis (RDA). These markers produce amplification products with genomic DNA from allotetraploid cotton Gossypium hirsutum. One of the markers is a polymorphic amplified restriction fragment (PARF) - a sequence found in both A and D genomes but differently flanked by restriction sites. Results of phylogenetic analysis of the PARF sequences from diploid cottons and from allotetraploid G. hirsutum agree with a previous observation of the interlocus concerted evolution (sequences corresponding to A and D genomes are homogenized to a D genome-type sequence). Our study shows how RDA can be used to develop genome-specific markers that can be used to study molecular evolution of allopolyploids.

DNA, Plant↗

An evolutionary approach reveals a high protein-coding capacity of the human genome.

We developed a new evolutionary method for identifying exons from genomic sequences and found 19000 potential coding exons that are absent from all existing annotations of the human genome. Of these, 13700 satisfied very stringent criteria and can with confidence be considered as novel exons. Evidently, a large number of new human genes can be identified using evolutionary approaches.

Animals↗

Detection of gene duplications and block duplications in eukaryotic genomes.

Several eukaryotic genomes have been completely sequenced and this provides an opportunity to investigate the extent and characteristics (e.g., single gene duplication, block duplication, etc.) of gene duplication in a genome. Detecting duplicate genes in a genome, however, is not a simple problem because of several complications such as domain shuffling, the existence of isoforms derived from alternative splicing, and annotational errors in the databases. We describe a method for overcoming these difficulties and the extents of gene duplication in the genomes of Drosophila melanogaster, Caenorhabditis elegans, and yeast inferred from this method. We also describe a method for detecting block duplications in a genome. Application of this method showed that block duplication is a common phenomenon in both yeast and nematode. The patterns of block duplication in the two species are, however, markedly different. Yeast shows much more extensive block duplication than nematode, with some chromosomes having more than 40% of the duplications derived from block duplications. Moreover, in yeast the majority of block duplications occurred between chromosomes, while in nematode most block duplications occurred within chromosomes.

Animals↗

The K(A)/K(S) ratio test for assessing the protein-coding potential of genomic regions: an empirical and simulation study.

Comparative genomics is a simple, powerful way to increase the accuracy of gene prediction. In this study, we show the utility of a simple test for the identification of protein-coding exons using human/mouse sequence comparisons. The test takes advantage of the fact that in the vast majority of coding regions, synonymous substitutions (K(S)) occur much more frequently than nonsynonymous ones (K(A)) and uses the K(A)/K(S) ratio as the criterion. We show the following: (1) most of the human and mouse exons are sufficiently long and have a suitable degree of sequence divergence for the test to perform reliably; (2) the test is suited for the identification of long exons and single exon genes, which are difficult to predict by current methods; (3) the test has a false-negative rate, lower than most of current gene prediction methods and a false-positive rate lower than all current methods; (4) the test has been automated and can be used in combination with other existing gene-prediction methods.

Animals↗

Signatures of domain shuffling in the human genome.

To elucidate the role of exon shuffling in shaping the complexity of the human genome/proteome, we have systematically analyzed intron phase distributions in the coding sequence of human protein domains. We found that introns at the boundaries of domains show high excess of symmetrical phase combinations (i.e., 0-0, 1-1, and 2-2), whereas nonboundary introns show no excess symmetry. This suggests that exon shuffling has primarily involved rearrangement of structural and functional domains as a whole. Furthermore, we found that domains flanked by phase 1 introns have dramatically expanded in the human genome due to domain shuffling and that 1-1 symmetrical domains and domain families are nonrandomly distributed with respect to their age. The predominance and extracellular location of 1-1 symmetrical domains among domains specific to metazoans suggests that they are associated with the rise of multicellularity. On the other hand, 0-0 symmetrical domains tend to be over-represented among ancient protein domains that are shared between the eukaryotic and prokaryotic kingdoms, which is compatible with the suggestion of primordial domain shuffling in the progenote. To see whether the human data reflect general genomic patterns of metazoans, similar analyses were done for the nematode Caenorhabditis elegans. Although the C. elegans data generally concur with the human patterns, we identified fewer intron-bounded domains in this organism, consistent with the lower complexity of C. elegans genes. [The following individuals kindly provided reagents, samples, or unpublished information as indicated in the paper: Z. Gu and R. Stevens.]

Animals↗