Search PubMed⌕ Search

Biomedical subjects

Ravi Sachidanandam

Publications and source records attributed to Ravi Sachidanandam.

17 recordsLinked to original sources

Comprehensive splice-site analysis using comparative genomics.

We have collected over half a million splice sites from five species-Homo sapiens, Mus musculus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-and classified them into four subtypes: U2-type GT-AG and GC-AG and U12-type GT-AG and AT-AC. We have also found new examples of rare splice-site categories, such as U12-type introns without canonical borders, and U2-dependent AT-AC introns. The splice-site sequences and several tools to explore them are available on a public website (SpliceRack). For the U12-type introns, we find several features conserved across species, as well as a clustering of these introns on genes. Using the information content of the splice-site motifs, and the phylogenetic distance between them, we identify: (i) a higher degree of conservation in the exonic portion of the U2-type splice sites in more complex organisms; (ii) conservation of exonic nucleotides for U12-type splice sites; (iii) divergent evolution of C.elegans 3' splice sites (3'ss) and (iv) distinct evolutionary histories of 5' and 3'ss. Our study proves that the identification of broad patterns in naturally-occurring splice sites, through the analysis of genomic datasets, provides mechanistic and evolutionary insights into pre-mRNA splicing.

Animals↗

A germline-specific class of small RNAs binds mammalian Piwi proteins.

Small RNAs associate with Argonaute proteins and serve as sequence-specific guides to regulate messenger RNA stability, protein synthesis, chromatin organization and genome structure. In animals, Argonaute proteins segregate into two subfamilies. The Argonaute subfamily acts in RNA interference and in microRNA-mediated gene regulation using 21-22-nucleotide RNAs as guides. The Piwi subfamily is involved in germline-specific events such as germline stem cell maintenance and meiosis. However, neither the biochemical function of Piwi proteins nor the nature of their small RNA guides is known. Here we show that MIWI, a murine Piwi protein, binds a previously uncharacterized class of approximately 29-30-nucleotide RNAs that are highly abundant in testes. We have therefore named these Piwi-interacting RNAs (piRNAs). piRNAs show distinctive localization patterns in the genome, being predominantly grouped into 20-90-kilobase clusters, wherein long stretches of small RNAs are derived from only one strand. Similar piRNAs are also found in human and rat, with major clusters occurring in syntenic locations. Although their function must still be resolved, the abundance of piRNAs in germline cells and the male sterility of Miwi mutants suggest a role in gametogenesis.

Animals↗

Short/branched-chain acyl-CoA dehydrogenase deficiency due to an IVS3+3A>G mutation that causes exon skipping.

Short/branched-chain acyl-CoA dehydrogenase deficiency (SBCADD) is an autosomal recessive disorder of L: -isoleucine catabolism. Little is known about the clinical presentation associated with this enzyme defect, as it has been reported in only a limited number of patients. Because the presence of C5-carnitine in blood may indicate SBCADD, the disorder may be detected by MS/MS-based routine newborn screening. It is, therefore, important to gain more knowledge about the clinical presentation and the mutational spectrum of SBCADD. In the present study, we have studied two unrelated families with SBCADD, both with seizures and psychomotor delay as the main clinical features. One family illustrates the fact that affected individuals may also remain asymptomatic. In addition, the normal level of newborn blood spot C5-acylcarnitine in one patient underscores the fact that newborn screening by MS/MS currently lacks sensitivity in detecting SBCADD. Until now, seven mutations in the SBCAD gene have been reported, but only three have been tested experimentally. Here, we identify and characterize an IVS3+3A>G mutation (c.303+3A>G) in the SBCAD gene, and provide evidence that this mutation is disease-causing in both families. Using a minigene approach, we show that the IVS3+3A>G mutation causes exon 3 skipping, despite the fact that it does not appear to disrupt the consensus sequence of the 5' splice site. Based on these results and numerous literature examples, we suggest that this type of mutation (IVS+3A>G) induces missplicing only when in the context of non-consensus (weak) 5' splice sites. Statistical analysis of the sequences shows that the wild-type versions of 5' splice sites in which +3A>G mutations cause exon skipping and disease are weaker on average than a random set of 5' splice sites. This finding is relevant to the interpretation of the functional consequences of this type of mutation in other disease genes.

Alternative Splicing↗

Second-generation shRNA libraries covering the mouse and human genomes.

Loss-of-function phenotypes often hold the key to understanding the connections and biological functions of biochemical pathways. We and others previously constructed libraries of short hairpin RNAs that allow systematic analysis of RNA interference-induced phenotypes in mammalian cells. Here we report the construction and validation of second-generation short hairpin RNA expression libraries designed using an increased knowledge of RNA interference biochemistry. These constructs include silencing triggers designed to mimic a natural microRNA primary transcript, and each target sequence was selected on the basis of thermodynamic criteria for optimal small RNA performance. Biochemical and phenotypic assays indicate that the new libraries are substantially improved over first-generation reagents. We generated large-scale-arrayed, sequence-verified libraries comprising more than 140,000 second-generation short hairpin RNA expression plasmids, covering a substantial fraction of all predicted genes in the human and mouse genomes. These libraries are available to the scientific community.

Animals↗

GeneSeer: a sage for gene names and genomic resources.

BACKGROUND: Independent identification of genes in different organisms and assays has led to a multitude of names for each gene. This balkanization makes it difficult to use gene names to locate genomic resources, homologs in other species and relevant publications. METHODS: We solve the naming problem by collecting data from a variety of sources and building a name-translation database. We have also built a table of homologs across several model organisms: H. sapiens, M. musculus, R. norvegicus, D. melanogaster, C. elegans, S. cerevisiae, S. pombe and A. thaliana. This allows GeneSeer to draw phylogenetic trees and identify the closest homologs. This, in turn, allows the use of names from one species to identify homologous genes in another species. A website http://geneseer.cshl.org/ is connected to the database to allow user-friendly access to our tools and external genomic resources using familiar gene names. CONCLUSION: GeneSeer allows access to gene information through common names and can map sequences to names. GeneSeer also allows identification of homologs and paralogs for a given gene. A variety of genomic data such as sequences, SNPs, splice variants, expression patterns and others can be accessed through the GeneSeer interface. It is freely available over the web http://geneseer.cshl.org/ and can be incorporated in other tools through an http-based software interface described on the website. It is currently used as the search engine in the RNAi codex resource, which is a portal for short hairpin RNA (shRNA) gene-silencing constructs.

Alternative Splicing↗

GObar: a gene ontology based analysis and visualization tool for gene sets.

BACKGROUND: Microarray experiments, as well as other genomic analyses, often result in large gene sets containing up to several hundred genes. The biological significance of such sets of genes is, usually, not readily apparent. Identification of the functions of the genes in the set can help highlight features of interest. The Gene Ontology Consortium 1 has annotated genes in several model organisms using a controlled vocabulary of terms and placed the terms on a Gene Ontology (GO), which comprises three disjoint hierarchies for Molecular functions, Biological processes and Cellular locations. The annotations can be used to identify functions that are enriched in the set, but this analysis can be misleading since the underlying distribution of genes among various functions is not uniform. For example, a large number of genes in a set might be kinases just because the genome contains many kinases. RESULTS: We use the Gene Ontology hierarchy and the annotations to pick significant functions and pathways by comparing the distribution of functions in a given gene list against the distribution of all the genes in the genome, using the hypergeometric distribution to assign probabilities. GObar is a web-based visualizer that implements this algorithm. The public website for GObar 2 can analyse gene lists from the yeast (S. cervisiae), fly (D. Melanogaster), mouse (M. musculus) and human (H. sapiens) genomes. It also allows visualization of the GO tree, as well as placement of a single gene on the GO hierarchy. We analyse a gene list from a genomic study of pre-mRNA splicing to demonstrate the utility of GObar. CONCLUSION: GObar is freely available as a web-based tool at http://katahdin.cshl.org:9331/GO2 and can help analyze and visualize gene lists from genomic analyses.

Algorithms↗

High-density single-nucleotide polymorphism maps of the human genome.

Here we report a large, extensively characterized set of single-nucleotide polymorphisms (SNPs) covering the human genome. We determined the allele frequencies of 55,018 SNPs in African Americans, Asians (Japanese-Chinese), and European Americans as part of The SNP Consortium's Allele Frequency Project. A subset of 8333 SNPs was also characterized in Koreans. Because these SNPs were ascertained in the same way, the data set is particularly useful for modeling. Our results document that much genetic variation is shared among populations. For autosomes, some 44% of these SNPs have a minor allele frequency > or =10% in each population, and the average allele frequency differences between populations with different continental origins are less than 19%. However, the several percentage point allele frequency differences among the closely related Korean, Japanese, and Chinese populations suggest caution in using mixtures of well-established populations for case-control genetic studies of complex traits. We estimate that approximately 7% of these SNPs are private SNPs with minor allele frequencies <1%. A useful set of characterized SNPs with large allele frequency differences between populations (>60%) can be used for admixture studies. High-density maps of high-quality, characterized SNPs produced by this project are freely available.

Alleles↗

RNAi as a bioinformatics consumer.

RNAi has shown great potential for use as a tool for biological discovery, analysis and therapeutics. The involvement of the RNAi pathway in post-transcription silencing, transcriptional silencing and epigenetic silencing as well as its use as a tool for forward genetics and therapeutics throws up several bioinformatics challenges. This paper delineates several areas of research and reviews work that has already been done, the tools that are available and the challenges that lie ahead.

Computational Biology↗

Determinants of the inherent strength of human 5' splice sites.

We previously showed that the authentic 5' splice site (5'ss) of the first exon in the human beta-globin gene is intrinsically stronger than a cryptic 5'ss located 16 nucleotides upstream. Here we examined by mutational analysis the contribution of individual 5'ss nucleotides to discrimination between these two 5'ss. Based on the in vitro splicing efficiencies of a panel of 26 wild-type and mutant substrates in two separate 5'ss competition assays, we established a hierarchy of 5'ss and grouped them into three functional subclasses: strong, intermediate, and weak. Competition between two 5'ss from different subclasses always resulted in selection of the 5'ss that belongs to the stronger subclass. Moreover, each subclass has different characteristic features. Strong and intermediate 5'ss can be distinguished by their predicted free energy of base-pairing to the U1 snRNA 5' terminus (DeltaG). Whereas the extent of splicing via the strong 5'ss correlates well with the DeltaG, this is not the case for competition between intermediate 5'ss. Weak 5'ss were used only when the competing authentic 5'ss was inactivated by mutation. These results indicate that extensive complementarity to U1 snRNA exerts a dominant effect for 5'ss selection, but in the case of competing 5'ss with similarly modest complementarity to U1, the role of other 5'ss features is more prominent. This study reveals the importance of additional submotifs present in certain 5'ss sequences, whose characterization will be critical for understanding 5'ss selection in human genes.

Alternative Splicing↗

A resource for large-scale RNA-interference-based screens in mammals.

Gene silencing by RNA interference (RNAi) in mammalian cells using small interfering RNAs (siRNAs) and short hairpin RNAs (shRNAs) has become a valuable genetic tool. Here, we report the construction and application of a shRNA expression library targeting 9,610 human and 5,563 mouse genes. This library is presently composed of about 28,000 sequence-verified shRNA expression cassettes contained within multi-functional vectors, which permit shRNA cassettes to be packaged in retroviruses, tracked in mixed cell populations by means of DNA 'bar codes', and shuttled to customized vectors by bacterial mating. In order to validate the library, we used a genetic screen designed to report defects in human proteasome function. Our results suggest that our large-scale RNAi library can be used in specific, genetic applications in mammals, and will become a valuable resource for gene analysis and discovery.

Animals↗

From silencing to gene expression: real-time analysis in single cells.

We have developed an inducible system to visualize gene expression at the levels of DNA, RNA and protein in living cells. The system is composed of a 200 copy transgene array integrated into a euchromatic region of chromosome 1 in human U2OS cells. The condensed array is heterochromatic as it is associated with HP1, histone H3 methylated at lysine 9, and several histone methyltransferases. Upon transcriptional induction, HP1alpha is depleted from the locus and the histone variant H3.3 is deposited suggesting that histone exchange is a mechanism through which heterochromatin is transformed into a transcriptionally active state. RNA levels at the transcription site increase immediately after the induction of transcription and the rate of synthesis slows over time. Using this system, we are able to correlate changes in chromatin structure with the progression of transcriptional activation allowing us to obtain a real-time integrative view of gene expression.

Animals↗

Short hairpin activated gene silencing in mammalian cells.

RNA interference (RNAi) is now a popular method for silencing gene expression in a variety of systems. RNAi methods use double-stranded RNAs (dsRNAs) to target complementary RNAs for destruction. In mammalian systems, very short dsRNAs (22-25 bp) such as short interfering RNAs (siRNAs) or short hairpin RNAs (shRNAs) are used to avoid endogenous nonspecific antiviral responses that target longer dsRNAs. siRNAs elicit a transient silencing response, while shRNAs can be expressed continuously to establish stable gene silencing. shRNAs can be introduced into cells and animals using a variety of standard vectors as well as retroviral or lentiviral expression systems. This chapter describes the design, construction, validation, and use of shRNAs for silencing genes. We report our results from testing a variety of shRNA design features and shRNA expression vectors. We also provide methods that use shRNAs to permit different levels of gene expression. Additionally, we discuss some aspects important for constructing an information pipeline to support development of a large shRNA library.

Animals↗

Intrinsic differences between authentic and cryptic 5' splice sites.

Cryptic splice sites are used only when use of a natural splice site is disrupted by mutation. To determine the features that distinguish authentic from cryptic 5' splice sites (5'ss), we systematically analyzed a set of 76 cryptic 5'ss derived from 46 human genes. These cryptic 5'ss have a similar frequency distribution in exons and introns, and are usually located close to the authentic 5'ss. Statistical analysis of the strengths of the 5'ss using the Shapiro and Senapathy matrix revealed that authentic 5'ss have significantly higher score values than cryptic 5'ss, which in turn have higher values than the mutant ones. beta-Globin provides an interesting exception to this rule, so we chose it for detailed experimental analysis in vitro. We found that the sequences of the beta-globin authentic and cryptic 5'ss, but not their surrounding context, determine the correct 5'ss choice, although their respective scores do not reflect this functional difference. Our analysis provides a statistical basis to explain the competitive advantage of authentic over cryptic 5'ss in most cases, and should facilitate the development of tools to reliably predict the effect of disease-associated 5'ss-disrupting mutations at the mRNA level.

Alternative Splicing↗

A 3.9-centimorgan-resolution human single-nucleotide polymorphism linkage map and screening set.

Recent advances in technologies for high-throughout single-nucleotide polymorphism (SNP)-based genotyping have improved efficiency and cost so that it is now becoming reasonable to consider the use of SNPs for genomewide linkage analysis. However, a suitable screening set of SNPs and a corresponding linkage map have yet to be described. The SNP maps described here fill this void and provide a resource for fast genome scanning for disease genes. We have evaluated 6,297 SNPs in a diversity panel composed of European Americans, African Americans, and Asians. The markers were assessed for assay robustness, suitable allele frequencies, and informativeness of multi-SNP clusters. Individuals from 56 Centre d'Etude du Polymorphisme Humain pedigrees, with >770 potentially informative meioses altogether, were genotyped with a subset of 2,988 SNPs, for map construction. Extensive genotyping-error analysis was performed, and the resulting SNP linkage map has an average map resolution of 3.9 cM, with map positions containing either a single SNP or several tightly linked SNPs. The order of markers on this map compares favorably with several other linkage and physical maps. We compared map distances between the SNP linkage map and the interpolated SNP linkage map constructed by the deCode Genetics group. We also evaluated cM/Mb distance ratios in females and males, along each chromosome, showing broadly defined regions of increased and decreased rates of recombination. Evaluations indicate that this SNP screening set is more informative than the Marshfield Clinic's commonly used microsatellite-based screening set.

Alleles↗