Search PubMedSearch

SEARCH · Search PubMed

Results for “orthology”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation

Sequencing the orthologs of human autosomal forensic short tandem repeats provides individual- and species-level identification in African great apes.

BACKGROUND: Great apes are a global conservation concern, with anthropogenic pressures threatening their survival. Genetic analysis can be used to assess the effects of reduced population sizes and the effectiveness of conservation measures. In humans, autosomal short tandem repeats (aSTRs) are widely used in population genetics and for forensic individual identification and kinship testing. Traditionally, genotyping is length-based via capillary electrophoresis (CE), but there is an increasing move to direct analysis by massively parallel sequencing (MPS). An example is the ForenSeq DNA Signature Prep Kit, which amplifies multiple loci including 27 aSTRs, prior to sequencing via Illumina technology. Here we assess the applicability of this human-based kit in African great apes. We ask whether cross-species genotyping of the orthologs of these loci can provide both individual and (sub)species identification. RESULTS: The ForenSeq kit was used to amplify and sequence aSTRs in 52 individuals (14 chimpanzees; 4 bonobos; 16 western lowland, 6 eastern lowland, and 12 mountain gorillas). The orthologs of 24/27 human aSTRs amplified across species, and a core set of thirteen loci could be genotyped in all individuals. Genotypes were individually and (sub)species identifying. Both allelic diversity and the power to discriminate (sub)species were greater when considering STR sequences rather than allele lengths. Comparing human and African great-ape STR sequences with an orangutan outgroup showed general conservation of repeat types and allele size ranges. Variation in repeat array structures and a weak relationship with the known phylogeny suggests stochastic origins of mutations giving rise to diverse imperfect repeat arrays. Interruptions within long repeat arrays in African great apes do not appear to reduce allelic diversity. CONCLUSIONS: Orthologs of most human aSTRs in the ForenSeq DNA Signature Prep Kit can be analysed in African great apes. Primer redesign would reduce observed variability in amplification across some loci. MPS of the orthologs of human loci provides better resolution for both individual and (sub)species identification in great apes than standard CE-based approaches, and has the further advantage that there is no need to limit the number and size ranges of analysed loci.

Animals

Gene model for the ortholog of Glys in Drosophila ananassae.

Gene model for the ortholog of Glycogen synthase (Glys) in the May 2011 (Agencourt dana_caf1/DanaCAF1) Genome Assembly (GenBank Accession: GCA_000005115.1) of Drosophila ananassae. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Bioinformatics

Gene model for the ortholog of Pi3K21B in Drosophila eugracilis.

Gene model for the ortholog of Phosphatidylinositol 3-kinase 21B (Pi3K21B) in the D. eugracilis May 2021 (Stanford ASM1815383v1/DeugRefSeq2) Genome Assembly (GenBank Accession: GCF_018153835.1) of Drosophila eugracilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Journal Article

Gene model for the ortholog of foxo in Drosophila sechellia.

Gene model for the ortholog of forkhead box, sub-group O (foxo) in the May 2011 (Broad dsec_caf1/DsecCAF1) Genome Assembly (GenBank Accession: GCA_000005215.1) of Drosophila sechellia. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Bioinformatics

Gene Model for the ortholog of Ilp2 in Drosophila ananassae.

Gene model for the ortholog of Insulin-like peptide 2 ( Ilp2 ) in the D. ananassae May 2011 (Agencourt dana_caf1/DanaCAF1) Genome Assembly (GenBank Accession: GCA_000005115.1 ) of Drosophila ananassae . This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Journal Article

Gene model for the ortholog of Pi3K21B in Drosophila ananassae.

Gene model for the ortholog of Phosphatidylinositol 3-kinase 21B ( Pi3K21B ) in the May 2011 (Agencourt dana_caf1/DanaCAF1) Genome Assembly (GenBank Accession: GCA_000005115.1 ) of Drosophila ananassae . This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Journal Article

Gene model for the ortholog of dock in Drosophila ananassae.

Gene model for the ortholog of dreadlocks ( dock ) in the May 2011 (Agencourt dana_caf1/DanaCAF1) Genome Assembly (GenBank Accession: GCA_000005115.1 ) of Drosophila ananassae . This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.

Journal Article

Experimental strategy for characterization of novel TnpB orthologs.

TnpB proteins encoded in IS200/IS605 and IS607 mobile genetic elements are among the most widespread proteins in the microbial world. They function as RNA-guided DNA nucleases that play a critical role in transposon proliferation and are the predecessors of CRISPR-Cas12 effector proteins of the type V CRISPR-Cas family. Small size of TnpB nucleases makes them an attractive alternative for larger Cas9 and Cas12 proteins in genome editing applications. However, only a small fraction of TnpB nucleases characterized to date are active in human cells, highlighting the need to identify new TnpB variants that can function as genome editors. Here, we present an experimental pipeline for the characterization of TnpB proteins by combining in silico analysis with in vitro assays. To validate it we determined guide RNA and identified TAM for a set of TnpB orthologs. The proposed workflow can be employed for rapid screening and characterization of the huge TnpB protein family to identify novel TnpB variants that might expand the genome editing toolbox.

Humans

The Drosophila aryl hydrocarbon receptor ortholog, spineless, modulates survival and reproduction.

The aryl hydrocarbon receptor (AhR) is a highly conserved, ligand-activated transcription factor in mammals involved in multiple physiological processes, including development, xenobiotic detoxification, and potentially aging, and some of the most potent activators of AhR are tryptophan metabolites. AhR manipulation across species has been shown to have conflicting results on aging phenotypes that are often tissue specific. To expand our understanding of AhR and aging, we studied AhR affects survival and reproduction in Drosophila melanogaster, whose genome contains an ortholog of AhR, spineless (ss). Our findings indicate that ss-deficient flies have a shorter lifespan than wildtype flies but interestingly exhibit reduced mortality until approximately 40 days of age. Similarly, ss-deficient flies are more stress resistant than wildtype at young ages, but this reverses in later age. Negative lifespan-shortening effects of tryptophan metabolites were mitigated in ss-deficient flies, suggesting that the effects of these metabolites are ss-reliant, similar to AhR in mammals. Overall, our preliminary work demonstrates an evolutionarily conserved role for AhR in the aging process and increases our knowledge of the role of AhR/ss on aging phenotypes.

Animals

Comprehensive profiling of antibiotic resistance genes and functional clusters of orthologous groups annotation of gut microbiota in Indonesian Kedu chickens.

Antibiotic resistance is a growing global health concern, with poultry systems acting as important reservoirs of antibiotic resistance genes (ARGs). However, resistome and functional profiles of indigenous chickens raised under traditional systems remain underexplored. This study aimed to characterize the antibiotic resistome, virulence factor genes, and metabolic potential of gut microbiota in Indonesian Kedu chickens using a shotgun metagenomic approach. Digesta samples from five gastrointestinal segments of 21 healthy adult chickens were analyzed through high-throughput sequencing. ARGs were identified using the Comprehensive Antibiotic Resistance Database (CARD) and Antibiotic Resistance Genes Databases (ARDB), while virulence factors and functional genes were annotated using Virulence Factor Database (VFDB), Clusters of Orthologous Groups (COG), and Carbohydrate-Active EnZymes (CAZy) databases. Results revealed a diverse resistome dominated by multidrug resistance and efflux pump mechanisms, with prominent genes associated with fluoroquinolone, tetracycline, β-lactam, and glycopeptide resistance. The detection of clinically relevant ARGs suggests that genetic determinants associated with antimicrobial resistance are present in the gut microbiota of traditionally raised Kedu chickens, although metagenomic data alone cannot determine whether these genes are actively expressed or confer phenotypic resistance. Virulence factor analysis showed functions related to adherence, immune evasion, iron acquisition, quorum sensing, and efflux activity, reflecting strong microbial adaptability. Functional profiling demonstrated enrichment in translation, carbohydrate and amino acid metabolism, genome maintenance, and cell envelope biogenesis. Additionally, CAZyme analysis indicated a high capacity for complex polysaccharide degradation, supporting efficient utilization of fiber-rich traditional diets. In conclusion, this study provides a comprehensive metagenomic overview of antibiotic resistance and functional potential in Kedu chicken gut microbiota, emphasizing the importance of incorporating indigenous poultry into antimicrobial resistance surveillance within a One Health framework.

Antibiotic resistance genes

Nuclear single-copy orthologous genes as phylogenomic markers for resolving the closely related firefly genera Pteroptyx, Medeopteryx, and Trisinuata (Coleoptera: Lampyridae: Luciolinae).

Fireflies (Lampyridae) are bioluminescent beetles with broad ecological roles across temperate and tropical ecosystems, occupying diverse habitats including forests, wetlands, grasslands, mangroves, and riverine systems. The subfamily Luciolinae is primarily distributed across Asia and the Indo-Pacific. Phylogenetic relationships among three closely related Luciolinae genera - Medeopteryx, Pteroptyx, and Trisinuata - remain unresolved using mitochondrial genome data alone. This study used nuclear genome data to resolve relationships among these genera and identify a lighter-weight nuclear marker panel for expanding taxon sampling. Draft genomes were reconstructed for fifteen firefly species, eight from the focal genera, and analyzed with five published firefly genomes. Using BUSCO and OrthoFinder, 1,011 nuclear single-copy orthologs (SCOs) were identified for phylogenomic inference. Discordance between concatenation- and coalescence-based phylogenies indicated incomplete lineage sorting (ILS). The coalescence-based phylogeny recoveredPteroptyxas monophyletic and sister to a (Medeopteryx,Trisinuata) clade, with Trisinuata nested within a non-monophyletic Medeopteryx; however, quartet support at the base of Pteroptyx, particularly at Pt. valida, was low.Filtering for compositional homogeneity, clock-likeness, and species-tree concordance yielded 103 SCOs with a significantly higher proportion of parsimony-informative sites than non-selected loci, retaining the backbone topology with higher gene concordance support at scored clades, while ILS-driven discordance at Pt. valida persists - confirming that the reduced panel retains phylogenetic resolving power for future taxon sampling. These findings demonstrate a practical framework for using nuclear SCOs to resolve close phylogenetic relationships within Luciolinae. Future work should expand taxon sampling - especially forTrisinuata - alongside long-read assemblies, for a more robust phylogenomic framework.

Fireflies

Gene model for the ortholog of Ilp4 in Drosophila eugracilis.

Gene Model for Insulin-like peptide 4 (Ilp4) in the D. eugracilis (DeugGB2) assembly (GCA_000236325.2). The characterization of this ortholog was carried out as part of a larger, ongoing dataset designed to explore the evolution of the insulin/insulin-like growth factor signaling (IIS) pathway across the genus Drosophila, utilizing the Genomics Education Partnership gene annotation protocol within Course-based Undergraduate Research Experiences.

Bioinformatics

KINAID: an orthology-based kinase-substrate prediction and analysis tool for phosphoproteomics.

SUMMARY: Proteome-wide datasets of phosphorylated peptides, either measured in a condition of interest or in response to perturbations, are increasingly becoming available for model organisms across the evolutionary spectrum. We introduce KINAID (KINase Activity and Inference Dashboard), an interactive and extensible tool written in Dash/Plotly, that predicts kinase-substrate interactions, uncovers and displays kinases whose substrates are enriched amongst phosphorylated peptides, interactively illustrates kinase-substrate interactions, and clusters phosphopeptides targeted by similar kinases. KINAID is the first tool of its kind that can analyze data from not only Homo sapiens but also 10 additional model organisms (including Mus musculus, Danio rerio, Drosophila melanogaster, Caenorhabditis elegans, and Saccharomyces cerevisiae). We demonstrate KINAID's utility by applying it to recently published S. cerevisiae phosphoproteomics data. AVAILABILITY AND IMPLEMENTATION: Webserver is available at https://kinaid.princeton.edu; open-source python library is available at https://github.com/Singh-Lab/kinaid; archive is available at https://doi.org/10.24433/CO.8460107.v1.

Proteomics