Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans

Linkage group selection: rapid gene discovery in malaria parasites.

The identification of parasite genes controlling phenotypes such as drug resistance, virulence, immunogenicity, and transmission is vital to malaria research. Classical genetic methods have achieved these goals only rarely and with difficulty. We describe here a novel genetic method, Linkage Group Selection (LGS), which achieves rapid de novo location of genes encoding selectable phenotypes of malaria parasites. A phenotype-specific selection pressure is applied to the uncloned progeny of a genetic cross between two malaria parasites that differ in the relevant phenotype. Selected and unselected progeny are analyzed using genome-wide quantitative genetic markers. Markers of the "sensitive" parent, which are reduced after selection, are sequenced and located in genomic databases. They are expected to be closely linked to gene(s) determining the phenotype under selection. We have validated LGS with the rodent malaria parasite Plasmodium chabaudi chabaudi using a phenotype, pyrimethamine resistance, whose controlling gene, that encoding dihydrofolate reductase (dhfr), is known. We show that molecular markers closely linked to dhfr, and only those linked to this gene, were reduced or removed by pyrimethamine treatment in accordance with the expectations of LGS.

Animals

Altered neural electrophysiological properties in the anterior cingulate cortex in a mouse model of Prader-Willi syndrome.

Prader-Willi syndrome (PWS) is a neurodevelopmental genetic disease associated with multiple metabolic and behavioural abnormalities converging into a distinctive clinical phenotype characterized by insatiable appetite leading to hyperphagia and eventual morbid obesity. The PWS spectrum results from deficiencies in paternally imprinted chromosome 15q11-13 region clustering around non-coding RNA multiple-repeat gene Snord116. A PWS mouse model with paternal Snord116 deletion (Snord116del) revealed multiple expected behavioural traits but failed to reproduce obesity in experimental paradigms designed to uncover homeostatic hypothalamic mechanisms of hyperphagia, while the possibility for pathologic hedonic overdrive underlying hyperphagic behaviours was not studied. In Snord116del mice, we examined functional properties of pyramidal neurons (PyNs) in the anterior cingulate cortex (ACC), the brain area commonly associated with goal-oriented and choice-outcome processing, including the value assessment of food items. We found indications of higher dendritic complexity and stronger afferent excitatory connectivity compared to controls. A strong excitatory input into Snord116del PyNs was balanced by a more hyperpolarized resting membrane potential, rendering lower soma excitability, improved signal-to-noise discrimination and stronger low-pass filtering. The enhanced excitatory network-tuning ability originating from Snord116 deficiency may explain the previously reported better performance of Snord116del over wild-type mice in working-for-food behavioural tests, whereas in humans it might entail exaggerated reward-seeking behaviour since early childhood when food is the main attractant. Our analysis of previously published genomic databases revealed candidate genes responsible for the abnormal functional neuronal phenotype caused by Snord116 deletion, including K+ and Na+ voltage-dependent ion channels, protein kinases, phosphatases and components of the mechanistic target of rapamycin (mTOR) intracellular signalling pathway. KEY POINTS: Altered biophysical characteristics and parameters of neuronal connectivity in pyramidal neurons in the anterior cingulate cortex (ACC) in Snord116 deletion mice. Alterations include augmented afferent synaptic input, altered resting state and firing properties of ACC pyramidal neurons. Our findings uncover a possible mechanistic basis for altered ACC functionality in Prader-Willi syndrome.

Animals

FLT4 gene polymorphisms influence isolated ventricular septal defect predisposition in a Southwest China population.

BACKGROUND: Ventricular septal defect (VSD) is the most common congenital heart disease. Although a small number of genes associated with VSD have been found, the genetic factors of VSD remain unclear. In this study, we evaluated the association of 10 candidate single nucleotide polymorphisms (SNPs) with isolated VSD in a population from Southwest China. METHODS: Based on the results of 34 congenital heart disease whole-exome sequencing and 1000 Genomes databases, 10 candidate SNPs were selected. A total of 618 samples were collected from the population of Southwest China, including 285 VSD samples and 333 normal samples. Ten SNPs in the case group and the control group were identified by SNaPshot genotyping. The chi-square (&#x3c7;2) test was used to evaluate the relationship between VSD and each candidate SNP. The SNPs that had significant P value in the initial stage were further analysed using linkage disequilibrium, and haplotypes were assessed in 34 congenital heart disease whole-exome sequencing samples using Haploview software. The bins of SNPs that were in very strong linkage disequilibrium were further used to predict haplotypes by Arlequin software. ViennaRNA v2.5.1 predicted the haplotype mRNA secondary structure. We evaluated the correlation between mRNA secondary structure changes and ventricular septal defects. RESULTS: The &#x3c7;2 results showed that the allele frequency of FLT4 rs383985 (P&#x2009;=&#x2009;0.040) was different between the control group and the case group (P&#x2009;<&#x2009;0.05). FLT4 rs3736061 (r2&#x2009;=&#x2009;1), rs3736062 (r2&#x2009;=&#x2009;0.84), rs3736063 (r2&#x2009;=&#x2009;0.84) and FLT4 rs383985 were in high linkage disequilibrium (r2&#x2009;>&#x2009;0.8). Among them, rs3736061 and rs3736062 SNPs in the FLT4 gene led to synonymous variations of amino acids, but predicting the secondary structure of mRNA might change the secondary structure of mRNA and reduce the free energy. CONCLUSIONS: These findings suggest a possible molecular pathogenesis associated with isolated VSD, which warrants investigation in future studies.

Child

Identification of novel cytoskeleton protein involved in spermatogenic cells and sertoli cells of non-obstructive azoospermia based on microarray and bioinformatics analysis.

BACKGROUND: During mammalian spermatogenesis, the cytoskeleton system plays a significant role in morphological changes. Male infertility such as non-obstructive azoospermia (NOA) might be explained by studies of the cytoskeletal system during spermatogenesis. METHODS: The cytoskeleton, scaffold, and actin-binding genes were analyzed by microarray and bioinformatics (771 spermatogenic cellsgenes and 774 Sertoli cell genes). To validate these findings, we cross-referenced our results with data from a single-cell genomics database. RESULTS: In the microarray analyses of three human cases with different NOA spermatogenic cells, the expression of TBL3, MAGEA8, KRTAP3-2, KRT35, VCAN, MYO19, FBLN2, SH3RF1, ACTR3B, STRC, THBS4, and CTNND2 were upregulated, while expression of NTN1, ITGA1, GJB1, CAPZA1, SEPTIN8, and GOLGA6L6 were downregulated. There was an increase in KIRREL3, TTLL9, GJA1, ASB1, and RGPD5 expression in the Sertoli cells of three human cases with NOA, whereas expression of DES, EPB41L2, KCTD13, KLHL8, TRIOBP, ECM2, DVL3, ARMC10, KIF23, SNX4, KLHL12, PACSIN2, ANLN, WDR90, STMN1, CYTSA, and LTBP3 were downregulated. A combined analysis of Gene Ontology (GO) and STRING, were used to predict proteins' molecular interactions and then to recognize master pathways. Functional enrichment analysis showed that the biological process (BP) mitotic cytokinesis, cytoskeleton-dependent cytokinesis, and positive regulation of cell-substrate adhesion were significantly associated with differentially expressed genes (DEGs) in spermatogenic cells. Moleculare function (MF) of DEGs that were up/down regulated, it was found that tubulin bindings, gap junction channels, and tripeptide transmembrane transport were more significant in our analysis. An analysis of GO enrichment findings of Sertoli cells showed BP and MF to be common DEGs. Cell-cell junction assembly, cell-matrix adhesion, and regulation of SNARE complex assembly were significantly correlated with common DEGs for BP. In the study of MF, U3 snoRNA binding, and cadherin binding were significantly associated with common DEGs. CONCLUSION: Our analysis, leveraging single-cell data, substantiated our findings, demonstrating significant alterations in gene expression patterns.

Male

An Exosomal miRNA Biomarker for the Detection of Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) remains a difficult tumor to diagnose and treat. To date, PDAC lacks routine screening with no markers available for early detection. Exosomes are 40-150 nm-sized extracellular vesicles that contain DNA, RNA, and proteins. These exosomes are released by all cell types into circulation and thus can be harvested from patient body fluids, thereby facilitating a non-invasive method for PDAC detection. A bioinformatics analysis was conducted utilizing publicly available miRNA pancreatic cancer expression and genome databases. Through this analysis, we identified 18 miRNA with strong potential for PDAC detection. From this analysis, 10 (MIR31, MIR93, MIR133A1, MIR210, MIR330, MIR339, MIR425, MIR429, MIR1208, and MIR3620) were chosen due to high copy number variation as well as their potential to differentiate patients with chronic pancreatitis, neoplasms, and PDAC. These 10 were examined for their mature miRNA expression patterns, giving rise to 18 mature miRs for further analysis. Exosomal RNA from cell culture media was analyzed via RTqPCR and seven mature miRs exhibited statistical significance (miR-31-5p, miR-31-3p, miR-210-3p, miR-339-5p, miR-425-5p, miR-425-3p, and miR-429). These identified biomarkers can potentially be used for early detection of PDAC.

Humans

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

A Case Report of Infantile Dopa-Responsive Dystonia Onset With Sleep Disorder Complicated With Autism Spectrum Disorder.

AIMS/BACKGROUND: Dopa-responsive dystonia (DRD) is a rare genetic disorder with complex and diverse clinical manifestations, resulting in a high rate of misdiagnosis. This case report describes an infantile case of DRD complicated by autism spectrum disorder (ASD), initially presenting with a sleep disorder. We aim to summarize its clinical manifestations, diagnostic process, treatment, and follow-up outcomes in order to improve clinical understanding of this disease. CASE PRESENTATION: A retrospective analysis was performed on a male infant who was treated at Jinhua Maternal and Child Health Care Hospital in 2020. The patient presented at one month of age with sleep disturbances, delayed motor development, and intermittent upward deviation of the eyes. Genetic testing identified two heterozygous pathogenic variants in the tyrosine hydroxylase (TH) gene. Among them, the c.738-2A>G variant was not recorded in the Exome Aggregation Consortium (ExAC), Genome Aggregation Database (gnomAD), or 1000 Genomes Asian population databases. During follow-up, the patient was also found to have comorbid ASD. RESULTS: Genetic testing confirmed biallelic TH mutations, establishing the diagnosis of infantile DRD. The patient exhibited marked clinical response to levodopa/benserazide, though dose titration was required with growth. CONCLUSION: For infants with unexplained sleep disorder accompanied by delayed motor development, genetic testing should be performed as early as possible to facilitate the identification of the root cause and implement timely treatment. In addition, close follow-up should be conducted to detect comorbid neurodevelopmental disorders.

Humans

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software

Comparison of whole-genome sequencing-based analysis methods for taxonomic classification of isolates unclassified by MALDI-TOF MS.

Taxonomic identification of clinical isolates is routinely achieved using matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS). If the species cannot be reliably identified, whole-genome sequencing can be applied. The aim of this study was to compare the results of approaches for taxonomic assignment for classification of isolates that are difficult to identify. Fifty-seven isolates were included in the study. The isolates were whole-genome sequenced and de novo assembled. Assembly-based classification was performed with the Genome Taxonomy Database Toolkit (GTDB-Tk), BLAST against 16S rRNA gene databases, the Type (Strain) Genome Server (TYGS), and ribosomal MLST (rMLST). Read-based classification was performed with MetaPhlAn4 and Kraken2. Thirty-two isolates were assigned to the same species with all four assembly-based classifiers, while the remaining 25 showed diverging assignments. When evaluating the results for the latter isolates, GTDB-Tk performed better than the other classifiers regarding which assignments were most likely correct. Of the read-based classifiers, MetaPhlAn4 performed better than Kraken2. Our evaluation identified GTDB-Tk to be the strongest tool for taxonomic assignment of isolates that are difficult to identify. Disagreements between classifiers are likely due to database limitations, wrongly assigned taxonomy, or unreliable 16S rRNA gene-based assignments.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Twelve Japanese patients with POLG-related disorders: Population-specific genetic differences of POLG variants in Japan and Europe.

BACKGROUND: POLG encodes mitochondrial DNA (mtDNA) polymerase &#x3b3;. Pathogenic POLG variants cause mitochondrial diseases, including progressive external ophthalmoplegia. POLG-related disorders are relatively common in Europe, possibly because of the high prevalence of carriers in the general population, but remain rare in Japan for unclear reasons. METHODS: We performed long-range PCR on mtDNA from skeletal muscle and/or peripheral blood from 3146 patients with suspected mitochondrial disease between 1993 and 2021. We selected 167 individuals with clinical features suggestive of POLG-related disorders for POLG gene analysis; all lacked pathogenic mtDNA point mutations, and most had multiple mtDNA deletions and/or a family history of mitochondrial disease. RESULTS: Among the 167 patients (median age: 52&#xa0;years, range: 0-83&#xa0;years, 11% pediatric cases), we identified 12 Japanese patients with POLG-related disorders and six POLG variants, including one novel variant. The six variants were p.Y955C, p.R943H, p.T599I, p.M299L, p.Y1210* (c.3626_3629dupGATA), and the novel variant p.F377S (c.1130T>C). Neither these six variants nor the 10 previously reported cases from Japan included the POLG variants that are more frequent in Europe. We also analyzed three population databases: two whole-genome sequencing databases covering 61,000 and 9850 Japanese individuals, respectively, and one global population database (gnomAD) covering 730,000 individuals worldwide. POLG variants that are more frequent in Europe were not detected in the Japanese databases or among East Asian individuals in gnomAD. CONCLUSIONS: Our findings suggest population-specific genetic differences in POLG between Japanese and European populations, explaining the lower frequency of POLG-related disorders in Japan.

CPEO

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation

Dental wastewater reveals a hidden reservoir of oral bacteriophage diversity.

Bacteriophages (phages) are being explored as alternatives or complements to antibiotics because of their ability to selectively kill bacterial pathogens. However, phages that infect many oral bacteria remain undiscovered. Here, we discovered that dental wastewater harbors previously underexplored phage diversity. Viral particles concentrated from dental wastewater displayed diverse morphologies, including abundant filamentous phage-like particles. Deep long-read metagenomic sequencing of concentrated viral particles generated 7.4 billion bases of sequence data and yielded 255 medium- to high-quality viral operational taxonomic units (vOTUs), including 46 predicted complete genomes. Comparison with large phage databases revealed that 63 of these 255 vOTUs had no detectable match, indicating that extensive sequencing of dental wastewater substantially expands the number of potential bacteriophages associated with the human oral microbiome. Host prediction linked many vOTUs to oral-associated bacterial taxa, including species with few or no previously reported phages, such as Porphyromonas gingivalis, Tannerella forsythia, and Candidatus Saccharibacteria. Functional annotation identified diverse genes associated with antiphage defense systems within a subset of vOTUs, suggesting that oral phages may contribute to the movement of genes encoding bacterial immune functions within the oral microbiome. Together, these findings expand the known oral phageome and show that dental wastewater contains a largely untapped diversity of phages.IMPORTANCEThe human oral cavity contains a diverse microbial community, but the bacteriophages (phages) that infect many oral bacteria remain poorly characterized. This gap limits our understanding of how phages shape oral microbial communities. Here, we show that dental wastewater is an underexplored source of oral phage diversity. Deep long-read metagenomic sequencing revealed 255 medium- to high-quality phage operational taxonomic units, many of which are not present in existing oral phage databases. These genomes include predicted phages of periodontal disease-associated bacteria and other oral taxa with few or no known phages. Dental wastewater therefore expands the known human oral phageome and reveals candidate phages linked to bacteria associated with oral health and disease.

Bacteriophages

PAHG: the database of human multi-gene families.

BACKGROUND: In the early vertebrate history, gene duplications, including single-gene, segmental-gene (SSD), and whole-genome duplication (WGD), formed multigene families. Despite efforts to classify metazoan multigene families hierarchically for evolutionary insight, a gap exists in accessible, curated resources for human/vertebrate multigene families. RESULTS: Addressing this, we present the Phylogenomic Analysis of Human Genome (PAHG) database. It focuses on curated multigene families in the human genome, particularly within four paralogons: HOX-bearing (Hsa:2/7/12/17), FGFR-bearing (Hsa:4/5/8/10), MHC-bearing (Hsa:1/6/9/19), and chromosomes 1/2/8/20. CONCLUSION: The current PAHG version details the phylogenetic history of 221 human multigene families (1247 gene members) with 15,231 protein sequences from diverse metazoans. It provides insights into gene duplication timings, co-duplication events, and their relationships with human genome syntenic organization. The PAHG database addresses the lack of accessible resources, offering valuable information on human/vertebrate multigene family evolution. Access the PAHG database at: https://www.pahgncb.com/ and http://pahg.qau.edu.pk/ . This resource enriches our understanding of vertebrate genetic evolution.

Humans

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (&#xb1;5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

A fast comparative genome browser for diverse bacteria and archaea.

Genome sequencing has revealed an incredible diversity of bacteria and archaea, but there are no fast and convenient tools for browsing across these genomes. It is cumbersome to view the prevalence of homologs for a protein of interest, or the gene neighborhoods of those homologs, across the diversity of the prokaryotes. We developed a web-based tool, fast.genomics, that uses two strategies to support fast browsing across the diversity of prokaryotes. First, the database of genomes is split up. The main database contains one representative from each of the 6,377 genera that have a high-quality genome, and additional databases for each taxonomic order contain up to 10 representatives of each species. Second, homologs of proteins of interest are identified quickly by using accelerated searches, usually in a few seconds. Once homologs are identified, fast.genomics can quickly show their prevalence across taxa, view their neighboring genes, or compare the prevalence of two different proteins. Fast.genomics is available at https://fast.genomics.lbl.gov.

Archaea

A python based automated computational framework to classify and comparative genomics analysis of the global diversity of chili leaf curl virus (ChiLCV) strains to understand virus host interactions.

Chili leaf curl virus (ChiLCV) is a Begomovirus chillicapsici that is one of the most devastating viruses impacted on the production of chili in the world, especially in South Asia. In the present study, we combined high-throughput computational genomics with experimental analysis of global diversity. A workflow was created using automated Python scripts to download, curate and process ChiLCV genomes from public database. About 410 complete ChiLCV genomes download from public databases. Using a phylogenetic approach, these isolates were subdivided into 34 strains, belonging to 10 major clades, showing significant genetic diversity. Geographic analysis revealed that Pakistan (207 isolates) and India (148 isolates) were the main sources of ChiLCV diversity and the remainder of the isolates were from Oman, Bangladesh, Iran, Saudi Arabia and Sri Lanka. Recombination was observed as a major evolutionary force as more than twenty recombination events were detected. Analysis of cis-regulatory elements showed a complex structure of the viral promoter, including multiple binding sites for transcription factors, hormone-response elements, light-responsive elements, and stress-responsive elements, indicating a high number of interactions between viral regulatory elements and host signaling pathways. Pangenome analysis showed the presence of a highly dynamic open pangenome made up of strain-specific orthologous groups (species-specific orthogroups). Experimental inoculation of chili plants was also carried out to assess the biological effects of infection, along with phytochemical, FTIR, HPLC, and qPCR analyses.

Begomovirus

Whole-genome sequencing identifies genetic diversity and adaptive signatures of hypoxia and ultraviolet radiation in Chinese chickens.

INTRODUCTION: Domestic chickens primarily descended from the wild red junglefowl, play a crucial role in global egg and meat production. China hosts diverse indigenous chicken populations that have adapted to various environmental conditions, including high-altitude with hypoxic and ultraviolet radiation stress. METHOD: We analyzed whole-genome sequences of 118 birds from five Indigenous Chinese chicken populations and 295 chicken genomes from publicly available databases to identify genomic diversity, admixture, and selection signatures of chickens adapted to high-altitude environments. Selection signatures were identified using nucleotide diversity (&#x3c0;), Tajima's D, XPEHH, and XP-CLR, selection scan methods. RESULTS: We observed a reduction in genetic diversity and historical declines in effective population size in high-altitude chicken, suggesting ongoing selection pressures shaping these populations. Selection scans identified nine genomic regions under strong positive selection, enriched for genes associated with hypoxia and ultraviolet radiation. Notably, five genes (TPK1, BAZ2B, MARCHF7, LLGL2, and RCAN3) were repeatedly detected across multiple selection signature analyses. RNA-seq analysis further confirmed the differential expression of these genes in the lung and heart tissues of chickens adapted to high and low altitudes, reinforcing their role in physiological adaptation to hypoxic environments. Altitude adaptation is driven by the selection of genes involved in oxygen metabolism, cellular stress response, and energy regulation. CONCLUSION: Our study provides compelling genetic evidence for differentiation between high and low and high-altitude Chinese chicken populations. These findings also ensure our understanding of local adaptation in poultry and establish a genomic framework for breeding strategies to improve environmental resilience to altitude-related stressors.

Animals