Search PubMedSearch

SEARCH · Search PubMed

Results for “Molecular Sequence Annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

LncRAnalyzer: a robust workflow for long non-coding RNA discovery using RNA-Seq.

Long non-coding RNA (lncRNA) is a major transcript category that lacks protein-coding capabilities, with relatively low abundance and complex expression patterns. Distinguishing lncRNAs from protein-coding genes is a complex process involving multiple filtering steps. We developed an automated pipeline named LncRAnalyzer featuring retrained models for 60 species. This workflow aims to reduce the likelihood of obtaining protein-coding or partial protein-coding transcripts during lncRNA identification by utilizing eight distinct approaches. We conducted a 10-fold cross-validation of the sorghum models and training sets with their standard ones and other approaches using real-life RNA-Seq datasets and known lncRNA and CDS sequences of sorghum. The results showed that the sorghum models and training sets were outperformed. The pipeline output comprises upset plots illustrating the number of lncRNA/NPCTs identified by the approaches, commonly identified lncRNA and their classes, NPCTs, and expression count tables. A feature-level comparison and benchmarking analysis of LncRAnalyzer with four existing pipelines, namely, LncPipe, LncEvo, lncRNA-Annotation, and Plant-LncPipe, demonstrated that LncRAnalyzer is more comprehensive, easier to implement, and accurate in lncRNA predictions. This workflow also ascertains lncRNA origins from various Transposable Elements (TEs) in plants using TE annotations from APTEdb [http://apte.cp.utfpr.edu.br/]. LncRAnalyzer is publicly available on GitLab [https://gitlab.com/nikhilshinde0909/LncRAnalyzer.git] for academic users.

RNA, Long Noncoding

A novel deep learning-driven framework for improving lncRNA comprehensive annotation with LncADeep 2.0.

MOTIVATION: Long non-coding RNAs (lncRNAs) have emerged as crucial players in diverse physiological and pathological processes, yet the biological mechanisms of the vast majority of lncRNAs remain elusive. To fill this gap, it is necessary to improve the accuracy of lncRNA identification and functional annotation. RESULTS: Here, we introduce LncADeep 2.0, an integrated deep learning framework designed to meet these needs. In the identification module, LncADeep 2.0 incorporated novel peptide features along with sequence and structural information, demonstrating superior performance over our previous LncADeep and other existing tools on both annotated transcripts from GENCODE and RNA-seq data. For functional annotation, LncADeep 2.0 leveraged lncRNA-centric interaction networks and gene ontology terms through the transfer learning strategy to achieve robust annotation performance with limited functional data. Compared to LncADeep, LncADeep 2.0 could accurately elucidate the general functions of given lncRNA sequences, predict tissue- or cell-type-specific functions from bulk and single-cell RNA-seq data, and establish connections between tumor-associated lncRNAs and genomic markers. Overall, LncADeep 2.0 stands out as an efficient and reliable tool for lncRNA identification and functional annotation across a wide spectrum of biological processes. AVAILABILITY AND IMPLEMENTATION: LncADeep 2.0 is available for use at https://github.com/Jefferson-Chou/LncADeep2 and https://doi.org/10.5281/zenodo.17164767.

RNA, Long Noncoding

Genomes of Conopholis americana and Epifagus virginiana: two holoparasitic plants (Orobanchaceae).

Conopholis americana (American cancer-root) and Epifagus virginiana (beechdrops) are sister genera of holoparasitic plants (Orobanchaceae) native to eastern North America, parasitizing oaks and American beech, respectively. Both have served as models for plastid genome reduction, yet no nuclear genomes exist for either genus or any New World holoparasitic Orobanchaceae. Here we present the first nuclear genome assemblies for both species using PacBio HiFi sequencing. The C. americana assembly totals 1.82 Gb and E. virginiana totals 440 Mb, representing an approximately 4-fold difference in genome size between these sister genera. We observed a BUSCO completeness of 79% to 80% in both species, which is typical of holoparasites. While gene prediction identified 33,889 genes in C. americana and 21,031 in E. virginiana, repeat annotation revealed that LTR retrotransposons account for 78% of the genome size difference. These assemblies reveal contrasting mechanisms of genome evolution in sister holoparasitic genera and provide foundational resources for comparative genomics of parasitic plants.

Genome, Plant

Expanding kinetoplastid genome annotation through protein structure comparison.

Kinetoplastids belong to the Discoba supergroup, an early divergent eukaryotic clade. Although the amount of genomic information on these parasites has grown substantially, assigning gene functions through traditional sequence-based homology methods remains challenging. Recently, significant advancements have been made in in-silico protein structure prediction and algorithms for rapid and precise large-scale protein structure comparisons. In this work, we developed a protein structure-based homology search pipeline (ASC, Annotation by Structural Comparisons) and applied it to transfer biological information to all kinetoplastid proteins available in TriTrypDB, the reference database for this lineage. Our pipeline enabled the assignment of structural similarity to a substantial portion of kinetoplastid proteins, improving current knowledge through annotation transfer. Additionally, we identified structural homologs for representatives of 6,700 uncharacterized proteins across 33 kinetoplastid species, proteins that could not be annotated using existing sequence-based tools and databases. As a result, this approach allowed us to infer potential biological information for a considerable number of kinetoplastid proteins. Among these, we identified structural homologs to ubiquitous eukaryotic proteins that are challenging to detect in kinetoplastid genomes through standard genome annotation pipelines. The results (KASC, Kinetoplastid Annotation by Structural Comparison) are openly accessible to the community at kasc.fcien.edu.uy through a user-friendly, gene-by-gene interface that enables visual inspection of the data.

Kinetoplastida

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

Chromosome-level genome assembly of an Arctic fish species pale eelpout (Lycodes pallidus).

Eelpouts (Zoarcidae) are known for their bipolar distributions and distinctive biogeographic histories. However, limited genomic data have hindered our understanding of their adaptive evolution. In this study, we present a thoroughly annotated chromosome-level genome assembly of pale eelpout (Lycodes pallidus) generated through the integration of Illumina, PacBio circular consensus, and Hi-C sequencing techniques. The final assembly spans 753.4 Mb, with its high quality confirmed by a scaffold N50 of 28.6 Mb and a Benchmarking Universal Single-Copy Ortholog (BUSCO) completeness of 99.3%. In comparison to other eelpouts and related fishes, the L. pallidus genome is larger and exhibits greater repetitive element content, accounting for approximately 45% of its total length. We annotated 21,419 protein-coding genes, a significant proportion of which are involved in signal transduction mechanisms and transcription. These findings provide valuable genetic resources for elucidating the evolutionary mechanisms underlying polar fish adaptation.

Animals

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms

Chromosome-level genome assembly of the large carpenter bee Xylocopa dejeanii Lepeletier, 1841 (Hymenoptera: Apidae).

Xylocopinae, a diverse bee subfamily comprising over 1,000 bee species, and also a major model system for studying the pollination and evolution of sociality. The lack of chromosome-level genome assembly resources for the Xylocopinae limits our research of their biology and evolution. Here, we provided the first pseudo-chromosomes genome assembly of the Xylocopa dejeanii combined PacBio CLR long reads, Illumina sequences, and Hi-C data. The final genome is 194.44 Mb located in 16 chromosomes. Our assembly includes 141 scaffolds, with a scaffold N50 length of 13.15 Mb. BUSCO analysis revealed 99.00% completeness. Genome annotation identified 28.27 Mb of repetitive elements, 10,970 protein-coding genes, and 432 ncRNAs. This high-quality X. dejeanii assembly advances our understanding of Xylocopinae genomics and provides new insights into bee evolution.

Animals

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

A high-quality chromosome-level genome assembly and annotation of the giant freshwater prawn (Macrobrachium rosenbergii).

The giant freshwater prawn, Macrobrachium rosenbergii, is native to Southeast Asia and is used in aquacultural practices worldwide. It is considered advantageous because of its rapid growth, high nutritional value, and economic benefits. As one of the three major freshwater aquaculture shrimp sources in China, a high-quality genome resource is of great significance for promoting the germplasm improvement of varieties. This study presents a high-quality chromosome-level genome assembly of M. rosenbergii that was generated by combining PacBio, MGI, and Hi-C reads. The assembled genome was 2.96 Gb in size, with a contig N50 of 0.64 Mb and a scaffold N50 of 55.76 Mb, which was positioned on 59 pseudo-chromosomes. The Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis for genome assembly reached 94.37%. In total, 27,111 protein-coding genes were identified, of which 25,470 were functionally annotated. These results provide a foundation for future research into adaptive evolution, genomics, and molecular breeding in M. rosenbergii.

Animals

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

Genome assembly and annotation of the parasitoid jewel wasp Nasonia oneida.

The jewel wasp, Nasonia (Hymenoptera: Pteromalidae), is a well-established model system for evolutionary genetics and host-microbial interactions. Here, we present the genome of N. oneida, a species lacking prior genomic characterization, using 10× Genomics linked-read (400× coverage), Illumina short-read (120× coverage), and transcriptome data (30× coverage). The assembled genome size is 267 Mb, comprising 4,675 scaffolds, with a scaffold N50 of 1 Mb and 98.40% Benchmarking Universal Single-Copy Orthologues (BUSCOs) completeness score. Annotation revealed 32.29% (86.46 Mb) of repetitive sequences and 14,221 protein-coding genes. Comparative genomics of N. oneida with 15 other hymenopteran species validated the presence of 5,939 gene families shared among them, including 3643 single-copy and 2296 multicopy gene families. This study provides the first de novo assembly of N. oneida, providing a significant addition to the growing repertoire of molecular tools for comparative genomics and functional studies to understand the evolution of closely related species as well as the evolution of parasitic wasps.

Animals

Chromosome-level Genome Assembly of the Halophytic Turfgrass Zoysia macrostachya.

Zoysia macrostachya Franch. & Sav. is a halophytic perennial turfgrass in the Poaceae family, commonly found in the coastal regions of Korea, Japan, and East Asia. Z. macrostachya thrives in high-salinity environments, making it an excellent model for studying abiotic stress resilience. In this study, we present a chromosome-level genome assembly of Z. macrostachya, constructed using Oxford Nanopore long reads, Illumina short reads, and Omni-C sequencing data. The assembly spans 329.78 Mb across 20 chromosomes, with a scaffold N50 of 19.24 Mb, and includes complete telomeric sequences at both ends. The assembly showed 97.8% complete BUSCOs, indicating high genome completeness. Repeat element and gene annotation identified 44.03% of the genome as repetitive elements and 33,474 protein-coding genes. The gene annotation showed 97.1% complete BUSCOs and 86.92% functionally characterized genes. Macrosynteny analysis highlighted highly collinear relationships with related species, providing a foundational understanding of the Z. macrostachya genomic structure. This high-quality genome serves as a valuable resource for advancing salinity tolerance research and improving the genetic diversity of Zoysia species.

Genome, Plant

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals