Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Analysis of 10,000 ESTs from lymphocytes of the cynomolgus monkey to improve our understanding of its immune system.

BACKGROUND: The cynomolgus monkey (Macaca fascicularis) is one of the most widely used surrogate animal models for an increasing number of human diseases and vaccines, especially immune-system-related ones. Towards a better understanding of the gene expression background upon its immunogenetics, we constructed a cDNA library from Epstein-Barr virus (EBV)-transformed B lymphocytes of a cynomolgus monkey and sequenced 10,000 randomly picked clones. RESULTS: After processing, 8,312 high-quality expressed sequence tags (ESTs) were generated and assembled into 3,728 unigenes. Annotations of these uniquely expressed transcripts demonstrated that out of the 2,524 open reading frame (ORF) positive unigenes (mitochondrial and ribosomal sequences were not included), 98.8% shared significant similarities (E-value less than 1e-10) with the NCBI nucleotide (nt) database, while only 67.7% (E-value less than 1e-5) did so with the NCBI non-redundant protein (nr) database. Further analysis revealed that 90.0% of the unigenes that shared no similarities to the nr database could be assigned to human chromosomes, in which 75 did not match significantly to any cynomolgus monkey and human ESTs. The mapping regions to known human genes on the human genome were described in detail. The protein family and domain analysis revealed that the first, second and fourth of the most abundantly expressed protein families were all assigned to immunoglobulin and major histocompatibility complex (MHC)-related proteins. The expression profiles of these genes were compared with that of homologous genes in human blood, lymph nodes and a RAMOS cell line, which demonstrated expression changes after transformation with EBV. The degree of sequence similarity of the MHC class I and II genes to the human reference sequences was evaluated. The results indicated that class I molecules showed weak amino acid identities (<90%), while class II showed slightly higher ones. CONCLUSION: These results indicated that the genes expressed in the cynomolgus monkey could be used to identify novel protein-coding genes and revise those incomplete or incorrect annotations in the human genome by comparative methods, since the old world monkeys and humans share high similarities at the molecular level, especially within coding regions. The identification of multiple genes involved in the immune response, their sequence variations to the human homologues, and their responses to EBV infection could provide useful information to improve our understanding of the cynomolgus monkey immune system.

5' Untranslated Regions↗

Utility of genome sequencing and group-enrichment to support splice variant interpretation in Marfan syndrome.

PURPOSE: To quantify the impact of noncanonical FBN1 splice site variants in undiagnosed Marfan syndrome (MFS), a connective tissue disorder associated with skeletal abnormalities and familial thoracic aortic aneurysm disease (FTAAD). METHODS: A systematic analysis of ultrarare FBN1 variants was performed using genome sequencing data from the 100,000 Genomes Project. Variants were annotated with SpliceAI and the significance of enrichment among individuals with FTAAD was assessed using Fisher's exact test. Experimental validation used RNA sequencing, reverse transcriptase polymerase chain reaction, minigene constructs, and replication analysis was with data from UK Biobank. RESULTS: Using aggregate data for 78,195 individuals, we identified 13,864 singleton single-nucleotide variants in FBN1 of which 21 were predicted to affect splicing (SpliceAI > 0.5). Incidence of candidate splice variants in individuals recruited with FTAAD (9/703) was significantly elevated compared with that seen in non-FTAAD participants (12/77,492; odds ratio = 84, P = 9.7 &#xd7; 10-14). Additional analysis uncovered a further 14 families harboring 11 different FBN1 splice variants. A total of 20 candidate splice variants in 23 families were identified, of which 70% lay beyond the &#xb1;8 splice regions. RNA testing confirmed the predicted splice aberration in 16 of 20 and for 9 of 20, pseudoexonization was the likely splicing anomaly. CONCLUSION: Our findings indicate that noncanonical splice variants may account for approximately 3% of families with undiagnosed FTAAD, highlighting the importance of incorporating analysis of introns and confirmatory RNA testing into genetic testing for Marfan syndrome.

Humans↗

Novel genes derived from noncoding DNA in Drosophila melanogaster are frequently X-linked and exhibit testis-biased expression.

Descriptions of recently evolved genes suggest several mechanisms of origin including exon shuffling, gene fission/fusion, retrotransposition, duplication-divergence, and lateral gene transfer, all of which involve recruitment of preexisting genes or genetic elements into new function. The importance of noncoding DNA in the origin of novel genes remains an open question. We used the well annotated genome of the genetic model system Drosophila melanogaster and genome sequences of related species to carry out a whole-genome search for new D. melanogaster genes that are derived from noncoding DNA. Here, we describe five such genes, four of which are X-linked. Our RT-PCR experiments show that all five putative novel genes are expressed predominantly in testes. These data support the idea that these novel genes are derived from ancestral noncoding sequence and that new, favored genes are likely to invade populations under selective pressures relating to male reproduction.

Animals↗

Phylogeny of Na+/Ca2+ exchanger (NCX) genes from genomic data identifies new gene duplications and a new family member in fish species.

The Na+/Ca2+ exchanger (NCX) is a member of the cation/Ca2+ antiporter (CaCA) family and plays a key role in maintaining cellular Ca2+ homeostasis in a variety of cell types. NCX is present in a diverse group of organisms and exhibits high overall identity across species. To date, three separate genes, i.e., NCX1, NCX2, and NCX3, have been identified in mammals. However, phylogenetic analysis of the exchanger has been hindered by the lack of nonmammalian NCX sequences. In this study, we expand and diversify the list of NCX sequences by identifying NCX homologs from whole-genome sequences accessible through the Ensembl Genome Browser. We identified and annotated 13 new NCX sequences, including 4 from zebrafish, 4 from Japanese pufferfish, 2 from chicken, and 1 each from honeybee, mosquito, and chimpanzee. Examination of NCX gene structure, together with construction of phylogenetic trees, provided novel insights into the molecular evolution of NCX and allowed us to more accurately annotate NCX gene names. For the first time, we report the existence of NCX2 and NCX3 in organisms other than mammals, yielding the hypothesis that two serial NCX gene duplications occurred around the time vertebrates and invertebrates diverged. In addition, we have found a putative new NCX protein, named NCX4, that is related to NCX1 but has been observed only in fish species genomes. These findings present a stronger foundation for our understanding of the molecular evolution of the NCX gene family and provide a framework for further NCX phylogenetic and molecular studies.

Amino Acid Sequence↗

Complete genome sequence of bacteriophage T5.

The 121,752-bp genome sequence of bacteriophage T5 was determined; the linear, double-stranded DNA is nicked in one of the strands and has large direct terminal repeats of 10,139 bp (8.3%) at both ends. The genome structure is consistently arranged according to its lytic life cycle. Of the 168 potential open reading frames (ORFs), 61 were annotated; these annotated ORFs are mainly enzymes involved in phage DNA replication, repair, and nucleotide metabolism. At least five endonucleases that believed to help inducing nicks in T5 genomic DNA, and a DNA ligase gene was found to be split into two separate ORFs. Analysis of T5 early promoters suggests a probable motif AAA{3, 4 T}nTTGCTT{17, 18 n}TATAATA{12, 13 W}{10 R} for strong promoters that may strengthen the step modification of host RNA polymerase, and thus control transcription of phage DNA. The distinct protein domain profile and a mosaic genome structure suggest an origin from the common genetic pool.

Bacteriophages↗

Transcriptional maps of 10 human chromosomes at 5-nucleotide resolution.

Sites of transcription of polyadenylated and nonpolyadenylated RNAs for 10 human chromosomes were mapped at 5-base pair resolution in eight cell lines. Unannotated, nonpolyadenylated transcripts comprise the major proportion of the transcriptional output of the human genome. Of all transcribed sequences, 19.4, 43.7, and 36.9% were observed to be polyadenylated, nonpolyadenylated, and bimorphic, respectively. Half of all transcribed sequences are found only in the nucleus and for the most part are unannotated. Overall, the transcribed portions of the human genome are predominantly composed of interlaced networks of both poly A+ and poly A- annotated transcripts and unannotated transcripts of unknown function. This organization has important implications for interpreting genotype-phenotype associations, regulation of gene expression, and the definition of a gene.

Cell Line↗

Genome-wide analysis of mRNA lengths in Saccharomyces cerevisiae.

BACKGROUND: Although the protein-coding sequences in the Saccharomyces cerevisiae genome have been studied and annotated extensively, much less is known about the extent and characteristics of the untranslated regions of yeast mRNAs. RESULTS: We developed a 'Virtual Northern' method, using DNA microarrays for genome-wide systematic analysis of mRNA lengths. We used this method to measure mRNAs corresponding to 84% of the annotated open reading frames (ORFs) in the S. cerevisiae genome, with high precision and accuracy (measurement errors +/- 6-7%). We found a close linear relationship between mRNA lengths and the lengths of known or predicted translated sequences; mRNAs were typically around 300 nucleotides longer than the translated sequences. Analysis of genes deviating from that relationship identified ORFs with annotation errors, ORFs that appear not to be bona fide genes, and potentially novel genes. Interestingly, we found that systematic differences in the total length of the untranslated sequences in mRNAs were related to the functions of the encoded proteins. CONCLUSIONS: The Virtual Northern method provides a practical and efficient method for genome-scale analysis of transcript lengths. Approximately 12-15% of the yeast genome is represented in untranslated sequences of mRNAs. A systematic relationship between the lengths of the untranslated regions in yeast mRNAs and the functions of the proteins they encode may point to an important regulatory role for these sequences.

Blotting, Northern↗

Characterization of the genomic organization of the region bordering the centromere of chromosome V of Podospora anserina by direct sequencing.

A Podospora anserina BAC library of 4800 clones has been constructed in the vector pBHYG allowing direct selection in fungi. Screening of the BAC collection for centromeric sequences of chromosome V allowed the recovery of clones localized on either sides of the centromere, but no BAC clone was found to contain the centromere. Seven BAC clones containing 322,195 and 156,244bp from either sides of the centromeric region were sequenced and annotated. One 5S rRNA gene, 5 tRNA genes, and 163 putative coding sequences (CDS) were identified. Among these, only six CDS seem specific to P. anserina. The gene density in the centromeric region is approximately one gene every 2.8kb. Extrapolation of this gene density to the whole genome of P. anserina suggests that the genome contains about 11,000 genes. Synteny analyses between P. anserina and Neurospora crassa show that co-linearity extends at the most to a few genes, suggesting rapid genome rearrangements between these two species.

Amino Acid Sequence↗

Genomic expansion and clustering of ZAD-containing C2H2 zinc-finger genes in Drosophila.

C2H2 zinc-finger proteins (ZFPs) constitute the largest family of nucleic acid binding factors in higher eukaryotes. In silico analysis identified a total of 326 putative ZFP genes in the Drosophila genome, corresponding to approximately 2.3% of the annotated genes. Approximately 29% of the Drosophila ZFPs are evolutionary conserved in humans and/or Caenorhabditis elegans. In addition, approximately 28% of the ZFPs contain an N-terminal zinc-finger-associated C4DM domain (ZAD) consisting of approximately 75 amino acid residues. The ZAD is restricted to ZFPs of dipteran and closely related insects. The evolutionary restriction, an expansion of ZAD-containing ZFP genes in the Drosophila genome and their clustering at few chromosomal sites are features reminiscent of vertebrate KRAB-ZFPs. ZADs are likely to represent protein-protein interaction domains. We propose that ZAD-containing ZFP genes participate in transcriptional regulation either directly or through site-specific modification and/or regulation of chromatin.

Amino Acid Sequence↗

Experiments in searching small proteins in unannotated large eukaryotic genomes.

There is growing interest to use mass spectrometry data to search genome sequences directly. Previous work by other authors demonstrated that this approach is able to correct and complement available genome annotations. We discuss the practical difficulty of searching large eukaryotic genomes with peptide ion trap tandem mass spectra of small proteins (<40 kDa). The challenging problem of automatically identifying peptides that span across exon/intron boundaries is explored for the first time by using experimental data. In a human genome search, we find that roughly 30% of the peptides are missed, due to various reasons, compared to a Swiss-Prot search. We show that this percentage is significantly reduced with improved parent mass accuracy. We finally provide several examples of predicted gene structures that could be improved by proteomics data, in particular by peptides spanning across exon/intron boundaries.

Adult↗

Nested genes in the human genome.

Here we studied one special type of gene, i.e., the nested gene, in the human genome. We collected 373 reliably annotated nested genes. Two-thirds of them were on the strand opposite that of their host gene. About 58% coding nested gene pairs were conserved in mouse and some were even maintained in chicken and fish, while nested pseudogenes were poorly conserved. Ka/Ks analysis revealed that nested genes were under strong selection, although they did not demonstrate greater conservation than other genes. With microarray data we observed that two partners of one nested pair seemed to be expressed reciprocally. A significant proportion of nested genes were tissue-specifically expressed. Gene ontology analysis demonstrated that quite a number of nested genes participated in cellular signal transduction. Based on these observations, we think that nested genes are a group of genes with important physiological functions.

Animals↗

Molecular cloning and functional expression of the first two specific insect myosuppressin receptors.

The Drosophila Genome Project database contains the sequences of two genes, CG8985 and CG13803, which are predicted to code for G protein-coupled receptors. We cloned the cDNAs corresponding to these genes and found that their gene structures had not been correctly annotated. We subsequently expressed the coding regions of the two corrected receptor genes in Chinese hamster ovary cells and found that each of them coded for a receptor that could be activated by low concentrations of Drosophila myosuppressin (EC50,4 x 10(-8) M). The insect myosuppressins are decapeptides that generally inhibit insect visceral muscles. Other tested Drosophila neuropeptides did not activate the two receptors. In addition to the two Drosophila myosuppressin receptors, we identified a sequence in the genomic database from the malaria mosquito Anopheles gambiae that also very likely codes for a myosuppressin receptor. To our knowledge, this paper is the first report on the molecular identification of specific insect myosuppressin receptors.

Amino Acid Sequence↗

Aplf/Dna2 variants drive chromosomal fission and accelerate speciation in zokors.

Chromosomal fissions and fusions are common, yet the molecular mechanisms and implications in speciation remain poorly understood. Here, we confirm a fission event in one zokor species through multiple-omics and functional analyses. We traced this event to a mutation in a splicing enhancer of the DNA repair gene Aplf in the fission-bearing species, which caused exon skipping and produced a truncated protein that disrupted DNA repair. An intronic deletion in Dna2, known to facilitate neo-telomere formation when knocked out, reduced gene activity. These variants collectively drove chromosomal fission in this zokor species. The newly formed chromosome became fixed due to carrying essential genes and strong selective pressure. While geographic isolation likely initiated the divergence of this species and the sister one, the fission event and associated decline at the chromosome level in gene flow probably exacerbated the speciation process. Our work elucidates the genetic basis of chromosomal fission and underscores its role in speciation dynamics.

Multiomics↗

The genome sequence of Caenorhabditis briggsae: a platform for comparative genomics.

The soil nematodes Caenorhabditis briggsae and Caenorhabditis elegans diverged from a common ancestor roughly 100 million years ago and yet are almost indistinguishable by eye. They have the same chromosome number and genome sizes, and they occupy the same ecological niche. To explore the basis for this striking conservation of structure and function, we have sequenced the C. briggsae genome to a high-quality draft stage and compared it to the finished C. elegans sequence. We predict approximately 19,500 protein-coding genes in the C. briggsae genome, roughly the same as in C. elegans. Of these, 12,200 have clear C. elegans orthologs, a further 6,500 have one or more clearly detectable C. elegans homologs, and approximately 800 C. briggsae genes have no detectable matches in C. elegans. Almost all of the noncoding RNAs (ncRNAs) known are shared between the two species. The two genomes exhibit extensive colinearity, and the rate of divergence appears to be higher in the chromosomal arms than in the centers. Operons, a distinctive feature of C. elegans, are highly conserved in C. briggsae, with the arrangement of genes being preserved in 96% of cases. The difference in size between the C. briggsae (estimated at approximately 104 Mbp) and C. elegans (100.3 Mbp) genomes is almost entirely due to repetitive sequence, which accounts for 22.4% of the C. briggsae genome in contrast to 16.5% of the C. elegans genome. Few, if any, repeat families are shared, suggesting that most were acquired after the two species diverged or are undergoing rapid evolution. Coclustering the C. elegans and C. briggsae proteins reveals 2,169 protein families of two or more members. Most of these are shared between the two species, but some appear to be expanding or contracting, and there seem to be as many as several hundred novel C. briggsae gene families. The C. briggsae draft sequence will greatly improve the annotation of the C. elegans genome. Based on similarity to C. briggsae, we found strong evidence for 1,300 new C. elegans genes. In addition, comparisons of the two genomes will help to understand the evolutionary forces that mold nematode genomes.

Animals↗

Using ESTs to improve the accuracy of de novo gene prediction.

BACKGROUND: ESTs are a tremendous resource for determining the exon-intron structures of genes, but even extensive EST sequencing tends to leave many exons and genes untouched. Gene prediction systems based exclusively on EST alignments miss these exons and genes, leading to poor sensitivity. De novo gene prediction systems, which ignore ESTs in favor of genomic sequence, can predict such "untouched" exons, but they are less accurate when predicting exons to which ESTs align. TWINSCAN is the most accurate de novo gene finder available for nematodes and N-SCAN is the most accurate for mammals, as measured by exact CDS gene prediction and exact exon prediction. RESULTS: TWINSCAN_EST is a new system that successfully combines EST alignments with TWINSCAN. On the whole C. elegans genome TWINSCAN_EST shows 14% improvement in sensitivity and 13% in specificity in predicting exact gene structures compared to TWINSCAN without EST alignments. Not only are the structures revealed by EST alignments predicted correctly, but these also constrain the predictions without alignments, improving their accuracy. For the human genome, we used the same approach with N-SCAN, creating N-SCAN_EST. On the whole genome, N-SCAN_EST produced a 6% improvement in sensitivity and 1% in specificity of exact gene structure predictions compared to N-SCAN. CONCLUSION: TWINSCAN_EST and N-SCAN_EST are more accurate than TWINSCAN and N-SCAN, while retaining their ability to discover novel genes to which no ESTs align. Thus, we recommend using the EST versions of these programs to annotate any genome for which EST information is available.TWINSCAN_EST and N-SCAN_EST are part of the TWINSCAN open source software package http://genes.cse.wustl.edu/distribution/download_TS.html.

Algorithms↗

JIGSAW: integration of multiple sources of evidence for gene prediction.

MOTIVATION: Computational gene finding systems play an important role in finding new human genes, although no systems are yet accurate enough to predict all or even most protein-coding regions perfectly. Ab initio programs can be augmented by evidence such as expression data or protein sequence homology, which improves their performance. The amount of such evidence continues to grow, but computational methods continue to have difficulty predicting genes when the evidence is conflicting or incomplete. Genome annotation pipelines collect a variety of types of evidence about gene structure and synthesize the results, which can then be refined further through manual, expert curation of gene models. RESULTS: JIGSAW is a new gene finding system designed to automate the process of predicting gene structure from multiple sources of evidence, with results that often match the performance of human curators. JIGSAW computes the relative weight of different lines of evidence using statistics generated from a training set, and then combines the evidence using dynamic programming. Our results show that JIGSAW's performance is superior to ab initio gene finding methods and to other pipelines such as Ensembl. Even without evidence from alignment to known genes, JIGSAW can substantially improve gene prediction accuracy as compared with existing methods. AVAILABILITY: JIGSAW is available as an open source software package at http://cbcb.umd.edu/software/jigsaw.

Algorithms↗

Faithful expression of a tagged Fugu WT1 protein from a genomic transgene in zebrafish: efficient splicing of pufferfish genes in zebrafish but not mice.

The teleost fish are widely used as model organisms in vertebrate biology. The compact genome of the pufferfish, Fugu rubripes, has proven a valuable tool in comparative genome analyses, aiding the annotation of mammalian genomes and the identification of conserved regulatory elements, whilst the zebrafish is particularly suited to genetic and developmental studies. We demonstrate that a pufferfish WT1 transgene can be expressed and spliced appropriately in transgenic zebrafish, contrasting with the situation in transgenic mice. By creating both transgenic mice and transgenic zebrafish with the same construct, we show that Fugu RNA is processed correctly in zebrafish but not in mice. Furthermore, we show for the first time that a Fugu genomic construct can produce protein in transgenic zebrafish: a full-length Fugu WT1 transgene with a C-terminal beta-galactosidase fusion is spliced and translated correctly in zebrafish, mimicking the expression of the endogenous WT1 gene. These data demonstrate that the zebrafish:Fugu system is a powerful and convenient tool for dissecting both vertebrate gene regulation and gene function in vivo.

Alternative Splicing↗

A complete survey of Trichoderma chitinases reveals three distinct subgroups of family 18 chitinases.

Genome-wide analysis of chitinase genes in the Hypocrea jecorina (anamorph: Trichoderma reesei) genome database revealed the presence of 18 ORFs encoding putative chitinases, all of them belonging to glycoside hydrolase family 18. Eleven of these encode yet undescribed chitinases. A systematic nomenclature for the H. jecorina chitinases is proposed, which designates the chitinases corresponding to their glycoside hydrolase family and numbers the isoenzymes according to their pI from Chi18-1 to Chi18-18. Phylogenetic analysis of H. jecorina chitinases, and those from other filamentous fungi, including hypothetical proteins of annotated fungal genome databases, showed that the fungal chitinases can be divided into three groups: groups A and B (corresponding to class V and III chitinases, respectively) also contained the so Trichoderma chitinases identified to date, whereas a novel group C comprises high molecular weight chitinases that have a domain structure similar to Kluyveromyces lactis killer toxins. Five chitinase genes, representing members of groups A-C, were cloned from the mycoparasitic species H. atroviridis (anamorph: T. atroviride). Transcription of chi18-10 (belonging to group C) and chi18-13 (belonging to a novel clade in group B) was triggered upon growth on Rhizoctonia solani cell walls, and during plate confrontation tests with the plant pathogen R. solani. Therefore, group C and the novel clade in group B may contain chitinases of potential relevance for the biocontrol properties of Trichoderma.

3' Untranslated Regions↗