Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Intron annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Identification of mammalian microRNA host genes and transcription units.

To derive a global perspective on the transcription of microRNAs (miRNAs) in mammals, we annotated the genomic position and context of this class of noncoding RNAs (ncRNAs) in the human and mouse genomes. Of the 232 known mammalian miRNAs, we found that 161 overlap with 123 defined transcription units (TUs). We identified miRNAs within introns of 90 protein-coding genes with a broad spectrum of molecular functions, and in both introns and exons of 66 mRNA-like noncoding RNAs (mlncRNAs). In addition, novel families of miRNAs based on host gene identity were identified. The transcription patterns of all miRNA host genes were curated from a variety of sources illustrating spatial, temporal, and physiological regulation of miRNA expression. These findings strongly suggest that miRNAs are transcribed in parallel with their host transcripts, and that the two different transcription classes of miRNAs ('exonic' and 'intronic') identified here may require slightly different mechanisms of biogenesis.

Base Sequence↗

Genomic exploration of the hemiascomycetous yeasts: 4. The genome of Saccharomyces cerevisiae revisited.

Since its completion more than 4 years ago, the sequence of Saccharomyces cerevisiae has been extensively used and studied. The original sequence has received a few corrections, and the identification of genes has been completed, thanks in particular to transcriptome analyses and to specialized studies on introns, tRNA genes, transposons or multigene families. In order to undertake the extensive comparative sequence analysis of this program, we have entirely revisited the S. cerevisiae sequence using the same criteria for all 16 chromosomes and taking into account publicly available annotations for genes and elements that cannot be predicted. Comparison with the other yeast species of this program indicates the existence of 50 novel genes in segments previously considered as 'intergenic' and suggests extensions for 26 of the previously annotated genes.

Ascomycota↗

Characterization of RFLP probe sequences for gene discovery and SSR development in Sorghum bicolor (L.) Moench.

In this study, we collected and analyzed DNA sequence data for 789 previously mapped RFLP probes from Sorghum bicolor (L.) Moench. DNA sequences, comprising 894 non-redundant contigs and end sequences, were searched against three GenBank databases, nucleotide (nt), protein (nr) and EST (dbEST), using BLAST algorithms. Matching ESTs were also searched against nt and nr. Translated DNA sequences were then searched against the conserved domain database (CDD) to determine if functional domains/motifs were congruent with the proteins identified in previous searches. More than half (500/894 or 56%) of the query sequences had significant matches in at least one of the GenBank searches. Overall, proteins identified for 148 sequences (17%) were consistent among all searches, of which 66 sequences (7%) contained congruent coding domains. The RFLP probe sequences were also evaluated for the presence of simple sequence repeats (SSRs) and 60 SSRs were developed and assayed in an array of sorghum germplasm comprising inbreds, landraces and wild relatives. Overall, these SSR loci had lower levels of polymorphism ( D = 0.46, averaged over 51 polymorphic loci) compared with sorghum SSRs that were isolated by library hybridization screens ( D = 0.69, averaged over 38 polymorphic loci). This result was probably due to the relatively small proportion of di-nucleotide repeat-containing markers (42% of the total SSR loci) obtained from the DNA sequence data. These di-nucleotide markers also contained shorter repeat motifs than those isolated from genomic libraries. Based on BLAST results, 24 SSRs (40%) were located within, or near, previously annotated or hypothetical genes. We determined the location of 19 of these SSRs relative to putative coding regions. In general, SSRs located in coding regions were less polymorphic ( D = 0.07, averaged over three loci) than those from gene flanking regions, UTRs and introns ( D = 0.49, averaged over 16 loci). The sequence information and SSR loci generated through this study will be valuable for application to sorghum genetics and improvement, including gene discovery, marker-assisted selection, diversity and pedigree analyses, comparative mapping and evolutionary genetic studies.

Journal Article↗

Bicistronic and fused monocistronic transcripts are derived from adjacent loci in the Arabidopsis genome.

Comparisons of full-length cDNAs and genomic DNAs available for Arabidopsis thaliana described here indicate that some adjacent loci are transcribed into extremely long RNAs spanning two annotated genes. Once expressed, some of these transcripts are post-transcriptionally spliced within their coding and intergenic sequences to generate bicistronic transcripts containing two complete open reading frames. Others are spliced to generate monocistronic transcripts coding for fusion proteins with sequences derived from both loci. RT-PCR of several P450 transcripts in this collection indicates that these extended transcripts exist side by side with shorter monocistronic transcripts derived from the individual loci in each pair. The existence of these unusual transcripts highlights variations in the processes of transcription and splicing that could not possibly have been predicted in the algorithms used for genome annotation and splice site predictions.

Arabidopsis↗

Disruption of the neuronal PAS3 gene in a family affected with schizophrenia.

Schizophrenia and its subtypes are part of a complex brain disorder with multiple postulated aetiologies. There is evidence that this common disease is genetically heterogeneous, with many loci involved. In this report, we describe a mother and daughter affected with schizophrenia, who are carriers of a t(9;14)(q34;q13) chromosome. By mapping on flow sorted aberrant chromosomes isolated from lymphoblast cell lines, both subjects were found to have a translocation breakpoint junction between the markers D14S730 and D14S70, a 683 kb interval on chromosome 14q13. This interval was found to contain the neuronal PAS3 gene (NPAS3), by annotating the genomic sequence for ESTs and performing RACE and cDNA library screenings. The NPAS3 gene was characterised with respect to the genomic structure, human expression profile, and protein cellular localisation to gain insight into gene function. The translocation breakpoint junction lies within the third intron of NPAS3, resulting in the disruption of the coding potential. The fact that the bHLH and PAS domains are disrupted from the remaining parts of the encoded protein suggests that the DNA binding and dimerisation functions of this protein are destroyed. The daughter (proband), who is more severely affected, has an additional microdeletion in the second intron of NPAS3. On chromosome 9q34, the translocation breakpoint junction was defined between D9S752 and D9S972 and no genes were found to be disrupted. We propose that haploinsufficiency of NPAS3 contributes to the cause of mental illness in this family.

ATP-Binding Cassette Transporters↗

Visualizing the genome: techniques for presenting human genome data and annotations.

BACKGROUND: In order to take full advantage of the newly available public human genome sequence data and associated annotations, biologists require visualization tools ("genome browsers") that can accommodate the high frequency of alternative splicing in human genes and other complexities. RESULTS: In this article, we describe visualization techniques for presenting human genomic sequence data and annotations in an interactive, graphical format. These techniques include: one-dimensional, semantic zooming to show sequence data alongside gene structures; color-coding exons to indicate frame of translation; adjustable, moveable tiers to permit easier inspection of a genomic scene; and display of protein annotations alongside gene structures to show how alternative splicing impacts protein structure and function. These techniques are illustrated using examples from two genome browser applications: the Neomorphic GeneViewer annotation tool and ProtAnnot, a prototype viewer which shows protein annotations in the context of genomic sequence. CONCLUSION: By presenting techniques for visualizing genomic data, we hope to provide interested software developers with a guide to what features are most likely to meet the needs of biologists as they seek to make sense of the rapidly expanding body of public genomic data and annotations.

Alternative Splicing↗

Interspecies conservation of gene order and intron-exon structure in a genomic locus of high gene density and complexity in Plasmodium.

A 13.6 kb contig of chromosome 5 of Plasmodium berghei, a rodent malaria parasite, has been sequenced and analysed for its coding potential. Assembly and comparison of this genomic locus with the orthologous locus on chromosome 10 of the human malaria Plasmodium falciparum revealed an unexpectedly high level of conservation of the gene organisation and complexity, only partially predicted by current gene-finder algorithms. Adjacent putative genes, transcribed from complementary strands, overlap in their untranslated regions, introns and exons, resulting in a tight clustering of both regulatory and coding sequences, which is unprecedented for genome organisation of PLASMODIUM: In total, six putative genes were identified, three of which are transcribed in gametocytes, the precursor cells of gametes. At least in the case of two multiple exon genes, alternative splicing and alternative transcription initiation sites contribute to a flexible use of the dense information content of this locus. The data of the small sample presented here indicate the value of a comparative approach for Plasmodium to elucidate structure, organisation and gene content of complex genomic loci and emphasise the need to integrate biological data of all Plasmodium species into the P.falciparum genome database and associated projects such as PlasmodB to further improve their annotation.

Alternative Splicing↗

[Analysis, identification and correction of some errors of model refseqs appeared in NCBI Human Gene Database by in silico cloning and experimental verification of novel human genes].

We found that human genome coding regions annotated by computers have different kinds of many errors in public domain through homologous BLAST of our cloned genes in non-redundant (nr) database, including insertions, deletions or mutations of one base pair or a segment in sequences at the cDNA level, or different permutation and combination of these errors. Basically, we use the three means for validating and identifying some errors of the model genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS: (I) Evaluating the support degree of human EST clustering and draft human genome BLAST. (2) Preparation of chromosomal mapping of our verified genes and analysis of genomic organization of the genes. All of the exon/intron boundaries should be consistent with the GT/AG rule, and consensuses surrounding the splice boundaries should be found as well. (3) Experimental verification by RT-PCR of the in silico cloning genes and further by cDNA sequencing. And then we use the three means as reference: (1) Web searching or in silico cloning of the genes of different species, especially mouse and rat homologous genes, and thus judging the gene existence by ontology. (2) By using the released genes in public domain as standard, which should be highly homologous to our verified genes, especially the released human genes appeared in NCBI GENOME ANNOTATION PROJECT REFSEQS, we try to clone each a highly homologous complete gene similar to the released genes in public domain according to the strategy we developed in this paper. If we can not get it, our verified gene may be correct and the released gene in public domain may be wrong. (3) To find more evidence, we verified our cloned genes by RT-PCR or hybrid technique. Here we list some errors we found from NCBI GENOME ANNOTATION PROJECT REFSEQs: (1) Insert a base in the ORF by mistake which causes the frame shift of the coding amino acid. In detail, abase in the ORF of a gene is a redundant insertion, which causes a reading frame shift in the translation of an alternative protein, such as LOC124919 is wrong form of C17 orf32 (with mouse and rat orthologs determined by us). (2) Put together by mistake (with force). This is a wrong assembly of non-relating cDNA segment, such as LOC147007 is wrong form of C17orf32. (3) Mistakenly insert a base or one section of cDNA in the ORF which causes it ending beforehand, only coding cDNA sequence of N-terminal amino acids, incomplete. For example, LOC123722 is wrong form of SPRYD1, and even the human hypothetical gene LOC126250 or PDCD5 is wrong form of our PDCD5 (TFAR19). (4) Incomplete, only coding cDNA sequence of C-terminal amino acids. For example, human LOC149076 and mouse LOC230761 are wrong form of our verified human ZNF362 and mouse Zfp362, respectively. (5) Incomplete, only coding one section of coding protein cDNA sequence of correct gene ORF, lacking N-terminal and C-terminal amino acids sequence, and at the same time, mistakenly anticipates the first non-initiation codon amino acid of the incomplete protein amino acid as the initiation codon, e.g. anticipating L as M. For example, LOC200084 is wrong form of ZNF362. (6) Mistakenly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However,these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone,and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. The articles published in the past did not clearly point out the existence of mistakes in the NCBI human gene mode reference sequence. At the Seventh International Human Genome Conference held in April 2002, we first published the researching result on this aspect in the communication form of Posterly insert a base or one section of cDNA in the ORF, wrongly causing unwanted termination codon before the insertion, so the coding protein lacks the first part of the amino acids. For example, the GenBank Acc. No. AL096883 ( LOCUS No. HS323M22B) is wrong form of an experimentally verified human NM_012263 with mouse ortholog of BC010510 determined. (7) It may regard the polluted genomic sequence as complete gene cDNA sequence and anticipate the so-called single exon gene, even the real one, only a small ORF in the very long single exon mRNA, while there really exists termination code in the same phase of the upper part of the ORF initiation code, no other characters accord with the gene's condition. For example, LOC91126 is wrong form of ZNF362. (8) The anticipated genes only have ORF which has no EST proofs on both terminal sides. Depending on this ORF, a complete gene cDNA with double support of EST and human genome (there are termination codes at the same phase of the upper part of ORF) which indicates the anticipated ORF reference sequence may be incorrect. For example, LOC164395 may be wrong form of novel human gene bankit4590055. (9) A similar but smaller protein-coding gene is anticipated in the range of the human genome sequence that has the support of EST experimental proof, so other new anticipated gene may be incorrect. For example, LOC167563 may be wrong form of CMYA5. However, these errors can be corrected or avoided by using our strategy. Here we give one example in detail: Comparision of the sequence SPRYD1 with human hypothetical gene LOC123722. The TAA bases in the position of 478-480 in LOC123722 cDNA is redundant, which causes a reading frame shift in the translation of an alternative protein. The redundancy of GTAAA of LOC123722 is not supported by our experimental clone, and is almost fully rejected by human EST alignment, and is shown as the next intron sequence by genomic GT/AG organization analysis. The verification of cDNA or genomic DNA sequence of SPRYD1 implies that LOC123722 has a wrong stop codon within its ORF because of the prediction program, thus being not complete cds. To sum up, by combining bioinformatics analyses with experimental verification, we have found that there are many errors of at least nine kinds appeared in NCBI GENOME ANNOTATION PROJECT REFSEQs through BLAST of our cloned genes in non-redundant database, and our strategy is helpful in correcting them, such as LOC14907, LOC200084 and LOC91126 (all of them should be ZNF362, but are three different kinds of wrong forms of ZNF362), three model reference sequences predicted from NCBI contig NT_004511 by automated computational analysis using gene prediction method, or such as LOC124919 and LOC147007 (both should be C17orf32, but are two different kinds of wrong forms of C17orf32), two model reference sequences predicted from NCBI contig NT_010808 by automated computational analysis using gene prediction method. Therefore, the correct identification and annotation of novel human genes may be still a heavy task, which can be finished within a long period of time. So human genome coding regions annotated by computer should be used with caution. (ABSTRACT TRUNCATED)

Amino Acid Sequence↗

HIV insertions within and proximal to host cell genes are a common finding in tissues containing high levels of HIV DNA and macrophage-associated p24 antigen expression.

HIV integration within host cell genomic DNA is a requisite step of the viral infection cycle. Yet, characteristics of the sites of provirus integration within the host genome remain obscure. The authors present evidence that in diseased tissues showing a high level of HIV DNA and macrophage-associated HIV p24 antigen expression from end stage forms of HIV disease, HIV-1 integration sites were favored within genes and transcriptionally active host cell genomic loci. Using an inverse PCR (IPCR) technique that identified dominant integrated forms of HIV, clonal IPCR products were isolated from AIDS dementia, AIDS lymphoma, and angioimmunoblastic lymphadenopathy tissues. Thirty of 34 disease-associated HIV-1 insertions were identified within annotated and hypothetical genes, an unexpected but highly nonrandom genetic coding region association (p <.026). The 1% sensitivity thresholds used for HIV IPCR suggested some form of selective expansion of cells containing these HIV proviruses. Consistent with this interpretation were the HIV-1 insertion sites identified within introns of genes that encoded for factors associated with signal transduction, apoptosis, and transcription regulation. In addition, HIV-1 proviruses were frequently found proximal to genes that encoded for receptor-associated, signal transduction-associated, transcription-associated, and translation-associated proteins. HIV-1 integration within host cell genomic DNA potentially represents a significant insertional mutagenic event. In certain cases, provirus insertions may mediate the dysregulation of specific gene expression events, providing mechanisms contributing to the pathogenesis associated with certain AIDS-related diseases.

Chromosomes, Human↗

Transcriptional analysis of a novel cluster of LY-6 family members in the human and mouse major histocompatibility complex: five genes with many splice forms.

Lymphocyte antigen-6 (LY-6) superfamily members are cysteine-rich, generally GPI-anchored cell surface proteins, which have definite or putative immune related roles. A cluster of five potential LY-6 superfamily members is located in the human and mouse major histocompatibility complex class III region. Comparative analysis of their genomic and cDNA sequences allowed us to carry out detailed annotations of these genes. We analyzed their mRNA expression patterns by RT-PCR performed on human and mouse cell line and tissue RNA. Sequence analysis of the transcripts revealed splice variants of all these genes in humans, and all but one in mouse. These splice forms retained introns or intron fragments, mainly generating premature stop codons, such that the only potentially functional mRNA was the predicted form. In some cases, the mis-spliced form was the most abundant form, suggesting a control mechanism for gene expression. Each gene showed mRNA expression differences between human and mouse.

Alternative Splicing↗

The Arabidopsis genome sequence as a tool for genome analysis in Brassicaceae. A comparison of the Arabidopsis and Capsella rubella genomes.

The annotated Arabidopsis genome sequence was exploited as a tool for carrying out comparative analyses of the Arabidopsis and Capsella rubella genomes. Comparison of a set of random, short C. rubella sequences with the corresponding sequences in Arabidopsis revealed that aligned protein-coding exon sequences differ from aligned intron or intergenic sequences in respect to the degree of sequence identity and the frequency of small insertions/deletions. Molecular-mapped markers and expressed sequence tags derived from Arabidopsis were used for genetic mapping in a population derived from an interspecific cross between Capsella grandiflora and C. rubella. The resulting eight Capsella linkage groups were compared to the sequence maps of the five Arabidopsis chromosomes. Fourteen colinear segments spanning approximately 85% of the Arabidopsis chromosome sequence maps and 92% of the Capsella genetic linkage map were detected. Several fusions and fissions of chromosomal segments as well as large inversions account for the observed arrangement of the 14 colinear blocks in the analyzed genomes. In addition, evidence for small-scale deviations from genome colinearity was found. Colinearity between the Arabidopsis and Capsella genomes is more pronounced than has been previously reported for comparisons between Arabidopsis and different Brassica species.

Arabidopsis↗

CREB binds to multiple loci on human chromosome 22.

The cyclic AMP-responsive element-binding protein (CREB) is an important transcription factor that can be activated by hormonal stimulation and regulates neuronal function and development. An unbiased, global analysis of where CREB binds has not been performed. We have mapped for the first time the binding distribution of CREB along an entire human chromosome. Chromatin immunoprecipitation of CREB-associated DNA and subsequent hybridization of the associated DNA to a genomic DNA microarray containing all of the nonrepetitive DNA of human chromosome 22 revealed 215 binding sites corresponding to 192 different loci and 100 annotated potential gene targets. We found binding near or within many genes involved in signal transduction and neuronal function. We also found that only a small fraction of CREB binding sites lay near well-defined 5' ends of genes; the majority of sites were found elsewhere, including introns and unannotated regions. Several of the latter lay near novel unannotated transcriptionally active regions. Few CREB targets were found near full-length cyclic AMP response element sites; the majority contained shorter versions or close matches to this sequence. Several of the CREB targets were altered in their expression by treatment with forskolin; interestingly, both induced and repressed genes were found. Our results provide novel molecular insights into how CREB mediates its functions in humans.

Binding Sites↗

Features of Arabidopsis genes and genome discovered using full-length cDNAs.

Arabidopsis is currently the reference genome for higher plants. A new, more detailed statistical analysis of Arabidopsis gene structure is presented including intron and exon lengths, intergenic distances, features of promoters, and variant 5'-ends of mRNAs transcribed from the same transcription unit. We also provide a statistical characterization of Arabidopsis transcripts in terms of their size, UTR lengths, 3'-end cleavage sites, splicing variants, and coding potential. These analyses were facilitated by scrutiny of our collection of sequenced full-length cDNAs and much larger collection of 5'-ESTs, together with another set of full-length cDNAs from Salk/Stanford/Plant Gene Expression Center/RIKEN. Examples of alternative splicing are observed for transcripts from 7% of the genes and many of these genes display multiple spliced isoforms. Most splicing variants lie in non-coding regions of the transcripts. Non-canonical splice sites constitute less than 1% of all splice sites. Genes with fewer than four introns display reduced average mRNA levels. Putative alternative transcription start sites were observed in 30% of highly expressed genes and in more than 50% of the genes with low expression. Transcription start sites correlate remarkably well with a CG skew peak in the DNA sequences. The intergenic distances vary considerably, those where genes are transcribed towards one another being significantly shorter. New transcripts, missing in the current TIGR genome annotation and ESTs that are non-coding, including those antisense to known genes, are derived and cataloged in the Supplementary Material. They identify 148 new loci in the Arabidopsis genome. The conclusions drawn provide a better understanding of the Arabidopsis genome and how the gene transcripts are processed. The results also allow better predictions to be made for, as yet, poorly defined genes and provide a reference for comparisons with other plant genomes whose complete sequences are currently being determined. Some comparisons with rice are included in this paper.

Alternative Splicing↗

Identification of a novel gene (HSN2) causing hereditary sensory and autonomic neuropathy type II through the Study of Canadian Genetic Isolates.

Hereditary sensory and autonomic neuropathy (HSAN) type II is an autosomal recessive disorder characterized by impairment of pain, temperature, and touch sensation owing to reduction or absence of peripheral sensory neurons. We identified two large pedigrees segregating the disorder in an isolated population living in Newfoundland and performed a 5-cM genome scan. Linkage analysis identified a locus mapping to 12p13.33 with a maximum LOD score of 8.4. Haplotype sharing defined a candidate interval of 1.06 Mb containing all or part of seven annotated genes, sequencing of which failed to detect causative mutations. Comparative genomics revealed a conserved ORF corresponding to a novel gene in which we found three different truncating mutations among five families including patients from rural Quebec and Nova Scotia. This gene, termed "HSN2," consists of a single exon located within intron 8 of the PRKWNK1 gene and is transcribed from the same strand. The HSN2 protein may play a role in the development and/or maintenance of peripheral sensory neurons or their supporting Schwann cells.

Amino Acid Sequence↗

MitoRes: a resource of nuclear-encoded mitochondrial genes and their products in Metazoa.

BACKGROUND: Mitochondria are sub-cellular organelles that have a central role in energy production and in other metabolic pathways of all eukaryotic respiring cells. In the last few years, with more and more genomes being sequenced, a huge amount of data has been generated providing an unprecedented opportunity to use the comparative analysis approach in studies of evolution and functional genomics with the aim of shedding light on molecular mechanisms regulating mitochondrial biogenesis and metabolism. In this context, the problem of the optimal extraction of representative datasets of genomic and proteomic data assumes a crucial importance. Specialised resources for nuclear-encoded mitochondria-related proteins already exist; however, no mitochondrial database is currently available with the same features of MitoRes, which is an update of the MitoNuc database extensively modified in its structure, data sources and graphical interface. It contains data on nuclear-encoded mitochondria-related products for any metazoan species for which this type of data is available and also provides comprehensive sequence datasets (gene, transcript and protein) as well as useful tools for their extraction and export. DESCRIPTION: MitoRes http://www2.ba.itb.cnr.it/MitoRes/ consolidates information from publicly external sources and automatically annotates them into a relational database. Additionally, it also clusters proteins on the basis of their sequence similarity and interconnects them with genomic data. The search engine and sequence management tools allow the query/retrieval of the database content and the extraction and export of sequences (gene, transcript, protein) and related sub-sequences (intron, exon, UTR, CDS, signal peptide and gene flanking regions) ready to be used for in silico analysis. CONCLUSION: The tool we describe here has been developed to support lab scientists and bioinformaticians alike in the characterization of molecular features and evolution of mitochondrial targeting sequences. The way it provides for the retrieval and extraction of sequences allows the user to overcome the obstacles encountered in the integrative use of different bioinformatic resources and the completeness of the sequence collection allows intra- and interspecies comparison at different biological levels (gene, transcript and protein).

Animals↗

Integrating alternative splicing detection into gene prediction.

BACKGROUND: Alternative splicing (AS) is now considered as a major actor in transcriptome/proteome diversity and it cannot be neglected in the annotation process of a new genome. Despite considerable progresses in term of accuracy in computational gene prediction, the ability to reliably predict AS variants when there is local experimental evidence of it remains an open challenge for gene finders. RESULTS: We have used a new integrative approach that allows to incorporate AS detection into ab initio gene prediction. This method relies on the analysis of genomically aligned transcript sequences (ESTs and/or cDNAs), and has been implemented in the dynamic programming algorithm of the graph-based gene finder EuGENE. Given a genomic sequence and a set of aligned transcripts, this new version identifies the set of transcripts carrying evidence of alternative splicing events, and provides, in addition to the classical optimal gene prediction, alternative optimal predictions (among those which are consistent with the AS events detected). This allows for multiple annotations of a single gene in a way such that each predicted variant is supported by a transcript evidence (but not necessarily with a full-length coverage). CONCLUSIONS: This automatic combination of experimental data analysis and ab initio gene finding offers an ideal integration of alternatively spliced gene prediction inside a single annotation pipeline.

Algorithms↗

Tackling non-canonical splicing in arrhythmogenic cardiomyopathy to reduce the uncertain significance variants burden.

BACKGROUND: Splice-altering variants (SAVs), particularly those outside canonical splice sites, are an underappreciated contributor to inherited cardiovascular diseases. In arrhythmogenic cardiomyopathy (ACM), these variants frequently remain classified as of uncertain significance (VUS) due to limited predictive power and lack of transcript-level evidence, constraining genetic yield and clinical management. Our study aimed to determine the functional impact of SAVs in ACM genes and refine their classification using ACMG/AMP and ClinGen SVI criteria. METHODS: SAVs identified in 200 ACM probands underwent SpliceAI prediction, GTEx cardiac exon-usage annotation, and functional assessment using pSPL3-based minigene assays. Aberrant transcripts were quantified using Percent Splicing Alteration (PSA). Segregation data and ACMG/AMP criteria refined by ClinGen SVI were applied to integrate functional and clinical evidence for classification. RESULTS: Aberrant splicing was confirmed in 9/20 variants (45%), including synonymous, missense, and non-canonical intronic changes. SpliceAI scores correlated strongly with PSA values (R&#xb2;=0.86). Case-control burden testing revealed significant enrichment of splice-altering variants in DSP, DSG2, DSC2 and FLNC. Integrating predictive algorithms with experimental validation and segregation analysis markedly enhances reclassification of 16/20 variants (80%). CONCLUSION: Splicing defects beyond canonical sites significantly shape ACM genetic landscape. Integrating predictive models with experimental validation clarifies uncertain variants bridging the gap between genomic uncertainty and clinical decision-making.

Humans↗

Composition-sensitive analysis of the human genome for regulatory signals.

Known transcription regulatory signals which generally act as transcription factor binding sites (TFs) differ significantly in their base composition. Therefore, their occurrence in a genome largely depends on the local base composition. In an attempt to initiate an all human genome analysis for the occurrence of potential TFs, we systematically analyzed the GC-content of distinct functional regions (e. g., upstream and downstream gene regions, exons, long and short introns, repetitive elements) and correlated the frequencies of potential binding sites of a representative set of TFs in these regions. For these analyses, we used the pattern collection of the TRANSFAC database on transcriptional regulation, the information about functionally relevant combinations of them from the database TRANSCompel, and our new resource, TRANSGenomeTM, which provides an overall annotation of the human genome with emphasis on its regulatory characteristics. We show that the occurrence of sequence patterns with regulatory potential may be supported by, but cannot be fully explained by either the GC content of a whole chromosome or its putative promoter regions, nor by the information content of the patterns. Several patterns, HNF-3, NFAT, and GC box, show a clear overrepresentation in all promoter groups as well as in all chromosomes. Other patterns, like E2F and CRE-BP1, are underrepresented in all promoter groups as well as in all chromosomes in comparison with random sequences. Simultaneously, both patterns are over-represented in promoters in comparison with repetitive elements. We define several structural characteristics of the proximal promoters that differentiate them from other functional genomic regions. Two well-known promoter elements, GC- and TATA-boxes, are statistically enriched in promoters in comparison with random sequences, repetitive elements and exons. Altogether, our findings provide insights into the macroheterogeneity amongst the individual chromosomes, into the microheterogeneity among different functional regions of individual chromosomes, contribute to further understanding of structural organization of gene regulatory regions, and give first hints on the development of regulatory features during evolution.

Animals↗