Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comparative genomic analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Comparative Genomic Analysis of Multidrug-Resistant Escherichia coli Across Poultry-Human-Environmental Interfaces.

The emergence of multidrug-resistant (MDR) Escherichia coli in poultry represents a critical One Health concern, particularly in developing countries. This study employed a comparative genomic approach to investigate the genomic characteristics, antimicrobial resistance (AMR) profiles, virulence determinants, of poultry-derived MDR E. coli isolates from Bangladesh. Whole-genome sequencing of three representative MDR isolates, identified with 83 globally diverse poultry, human, and environmental E. coli genomes. Pangenome analysis identified the characteristic open pangenome of E. coli, with core genes comprising only 4.6% of the combined dataset. Resistome analysis shown diverse AMR determinants, including blaCTX-M, blaTEM, sul, tet, and qnrS1, associated with antibiotic inactivation and efflux mechanisms. Virulence profiling revealed diverse genes involved in adhesion (fim, csg), iron acquisition (ent, fep, chu), motility, and secretion systems, with core virulence genes exhibiting > 90% sequence identity, whereas accessory virulence genes were more variable. Plasmid analysis demonstrated heterogeneous replicon types, predominantly IncF and Col plasmids, indicating their role in horizontal gene transfer. Jaccard similarity indices revealed moderate to high genetic overlap with global strains (~0.63 for virulence genes and ~0.55 for AMR profiles), suggesting shared evolutionary backgrounds. Phylogenomic and MLST identified all Bangladeshi isolates as ST457, clustering within a globally distributed clonal complex linked to ST10 and ST131 lineages. These findings suggest that the three Bangladeshi poultry-derived E. coli isolates are genetically related to globally circulating strains while harboring extensive resistance and virulence determinants, emphasizing poultry as an important reservoir of MDR pathogens and reinforcing the need for strengthened antimicrobial stewardship and genomic surveillance.

Animals↗

Single-read sequence tags of a limited number of genomic DNA fragments provide an inexpensive tool for comparative genome analysis.

Single-read sequences from both ends of 415 3-kb average size genomic DNA fragments of Candida albicans were compared with the complete sequence data of Saccharomyces cerevisiae. Comparison at the protein level, translated DNA against protein sequences, revealed 138 sequence tags with clear similarity to S. cerevisiae proteins or open reading frames. One case of synteny was found for the open reading frames of RAD16 and LYS2, which are adjacent to each other in S. cerevisiae and C. albicans.

Adenosine Triphosphatases↗

Functional and comparative genomic analysis of the piebald deletion region of mouse chromosome 14.

Several developmentally important genomic regions map within the piebald deletion complex on distal mouse chromosome 14. We have combined computational gene prediction and comparative sequence analysis to characterize an approximately 4.3-Mb segment of the piebald region to identify candidate genes for the phenotypes presented by homozygous deletion mice. As a result we have ordered 13 deletion breakpoints, integrated the sequence with markers from a bacterial artificial chromosome (BAC) physical map, and identified 16 known or predicted genes and >1500 conserved sequence elements (CSEs) across the region. The candidate genes identified include Phr1 (formerly Pam) and Spry2, which are mouse homologs of genes required for development in Drosophila melanogaster. Gene content, order, and position are highly conserved between mouse chromosome 14 and the orthologous region of human chromosome 13. Our studies combining computational gene prediction with genetic and comparative genomic analyses provide insight regarding the functional composition and organization of this defined chromosomal region.

Animals↗

Comparative genome analysis and pathway reconstruction.

Pathway reconstruction builds on genome and biochemical data with the aim of reconstructing higher level interactions between identified enzymes in a specific genome, in particular the different enzyme pathways (species or individual/patient). Metabolite flow in a pathway is analyzed by different tools, such as elementary mode analysis. This reveals key enzymes and pharmacological targets in the enzyme network. An overview of bioinformatic tools and algorithms for these tasks, application examples and recent results from these techniques are presented. Target selection, drug development and optimization can all be sped up using these approaches.

Animals↗

Comparative genomic analysis of Streptococcus parasuis and Streptococcus suis reveals mobile element-associated enrichment of antimicrobial resistance and lack of detectable same-MGE colocalization with virulence-associated genes within stable species boundaries.

Streptococcus suis is a major porcine pathogen and a zoonotic agent that causes meningitis and septicemia in humans. Streptococcus parasuis, a recently recognized close relative, remains poorly characterized with regard to its clinical significance and genomic features. In this study, we generated a single-contig closed genome assembly with genome-wide DNA methylation profiles for S. parasuis strain A1, isolated from a diseased pig in Xinjiang, China, and complemented in silico genomic predictions with isolate-level experimental validation of antimicrobial resistance (AMR) genotypes, virulence genotypes, and phenotypic susceptibility for this reference strain. Using this high-quality genome as a reference anchor, we performed comparative genomic analyses across 195 streptococcal genomes, comprising 15 S. parasuis and 180 S. suis strains, to distinguish genome-level co-occurrence of resistance and virulence determinants from their physical colocalization on the same mobile genetic element (MGE).Species boundaries remained clearly delineated at the genomic level, with a median interspecies average nucleotide identity (ANI) of approximately 86.0%, compared with intraspecies ANI medians of 97.5% for S. parasuis and 96.2% for S. suis. Pangenome analysis identified 12,693 gene clusters, of which 1086 were core clusters, and functional annotation revealed significant differences in accessory gene repertoires between the two species. Within this stable genomic framework, S. parasuis genomes carried a higher AMR gene burden; strain A1 harbored 10 AMR genes, multiple virulence-associated genes, three genomic islands, and eight prophage regions. For strain A1, PCR validation confirmed six AMR genes and six virulence genes, and disk diffusion testing demonstrated a multidrug-resistant phenotype consistent with the genotypic profile.Among 235 predicted mobile elements, 19 harbored AMR genes and seven carried Virulence Factor Database (VFDB) homologs, but none carried both categories simultaneously. This finding reflects a lack of detectable same-MGE colocalization under the applied annotation and assembly framework; it should not be interpreted as evidence of biological physical decoupling. Under a random-placement model, the expected number of co-carrying regions was only 0.57, and the probability of observing zero co-carrying regions was P = 0.55. This negative result should be interpreted with caution, given the limited number of cargo-bearing regions and the predominantly draft status of most genomes. Furthermore, the A1 genome contained multiple restriction-modification systems, showed depletion of several methylation motif families in mobile regions, and had limited CRISPR spacer matching evidence, suggesting prior exposure to the relevant sequence space. None of the genomes met our predefined criteria for whole-genome convergence.Collectively, our results support a model in which S. parasuis accumulates AMR-related genes in a modular fashion via mobile elements within stable species boundaries, with no detectable same-MGE colocalization of AMR and virulence determinants under our analytical pipeline. These findings imply that AMR surveillance strategies for this species should prioritize tracking mobile genetic elements rather than inferring wholesale genomic convergence toward S. suis.

Streptococcus suis↗

[Comparative genomics analysis of nitrate and nitrite respiration in gamma proteobacteria].

Nitrate and nitrite are preferred respiration oxidants during anaerobic conditions. In Escherichia coli such nitrate- and nitrite respiration is controlled by homologous transcriptional factors NarL and NarP. Although this system was intensively studied during the last two decades, the exact mechanisms of regulation and the structure of the NarL binding signals remained elusive. By the use of comparative genomics approach it was determined that most of the gammaproteobacteria contained only NarP protein. Regulog analysis revealed that whole structure of NarP regulons varied in different genomes and only regulation of nitrate and nitrite reduction system seemed to be highly conservative. Correlation between changes in the respiration system and the presence of the single regulatory system was shown. Conservative NarP binding sites upstream of fnr gene and genes for aerobic metabolism point to alteration in NarP role in respiration control during evolution. Thirty five new regulog members were determined and autoregulation of narQP operon in Vibrionaceae genomes was predicted.

Amino Acid Sequence↗

A comparative genomic analysis of left- and right-sided colon cancer using real-world data from the AACR project GENIE BPC dataset.

Left- and Right-sided colon cancers (LCC and RCC) are increasingly recognized as distinct clinicopathological and molecular subtypes with divergent prognoses and therapeutic responses. Leveraging a large, multi-institutional cohort from the AACR Project Genomics Evidence Neoplasia Information Exchange (GENIE) Biopharma Collaborative (BPC) (n = 750; LCC: 363 vs. RCC: 387), we conducted a comprehensive analysis of mutational profiles, tumor mutation burden (TMB), and survival outcomes. Our findings revealed a markedly higher TMB in RCC compared to LCC (6.65 &#xb1; 11.3 vs. 3.17 &#xb1; 4.35; adjusted P = 3.12&#xd7;10-32), suggesting greater genomic instability in RCC. After applying functional annotation filters (PolyPhen > 0.85, SIFT < 0.05), RCC tumors were significantly enriched for mutations in BRAF (23.1% vs. 6.7%), KMT2D (8.6% vs. 3.2%), and SMAD4 (13.1% vs. 7.3%), while TP53 mutations predominated in LCC (40.6% vs. 31.8%). Multivariate Cox regression analysis identified RCC as an independent predictor of poorer overall survival (OS) relative to LCC (HR: 1.30, 95% CI: 1.02-1.66, P = 0.033). Notably, KRAS mutations were associated with significantly worse OS in LCC (HR: 1.68, 95% CI: 1.06-2.70, P = 0.027), while BRAF mutations predicted adverse outcomes in RCC (HR: 1.58, 95% CI: 1.05-2.37, P = 0.028). These results underscore the prognostic value of tumor sidedness and specific genetic alterations in colon adenocarcinoma. Our study highlights the need for sidedness-specific molecular profiling to inform precision oncology strategies in colon cancer management.

BRAF↗

Comparative genome analysis identifies the vitamin D receptor gene as a direct target of p53-mediated transcriptional activation.

p53 is the most frequently mutated tumor suppressor gene in human neoplasia and encodes a transcriptional coactivator. Identification of p53 target genes is therefore key to understanding the role of p53 in tumorigenesis. To identify novel p53 target genes, we first used a comparative genomics approach to identify p53 binding sequences conserved in the human and mouse genome. We hypothesized that potential p53 binding sequences that are conserved are more likely to be functional. Using stringent filtering procedures, 32 genes were newly identified as putative p53 targets, and their responsiveness to p53 in human cancer cells was confirmed by reverse transcription-PCR and real-time PCR. Among them, we focused on the vitamin D receptor (VDR) gene because vitamin D3 has recently been used for chemoprevention of human tumors. VDR is induced by p53 as well as several other p53 family members, and analysis of chromatin immunoprecipitation showed that p53 protein binds to conserved intronic sequences of the VDR gene in vivo. Introduction of VDR into cells resulted in induction of several genes known to be p53 targets and suppression of colorectal cancer cell growth. In addition, p53 induced VDR target genes in a vitamin D3-dependent manner. Our in silico approach is a powerful method for identification of functional p53 binding sites and p53 target genes that are conserved among humans and other organisms and for further understanding the function of p53 in tumorigenesis.

Animals↗

Comparative genomic analysis of Vibrio cholerae: genes that correlate with cholera endemic and pandemic disease.

Historically, the first six recorded cholera pandemics occurred between 1817 and 1923 and were caused by Vibrio cholerae O1 serogroup strains of the classical biotype. Although strains of the El Tor biotype caused sporadic infections and cholera epidemics as early as 1910, it was not until 1961 that this biotype emerged to cause the 7th pandemic, eventually resulting in the global elimination of classical biotype strains as a cause of disease. The completed genome sequence of 7th pandemic El Tor O1 strain N16961 has provided an important tool to begin addressing questions about the evolution of V. cholerae as a human pathogen and environmental organism. To facilitate such studies, we constructed a V. cholerae genomic microarray that displays over 93% of the predicted genes of strain N16961 as spotted features. Hybridization of labeled genomic DNA from different strains to this microarray allowed us to compare the gene content of N16961 to that of other V. cholerae isolates. Surprisingly, the results reveal a high degree of conservation among the strains tested. However, genes unique to all pandemic strains as well as genes specific to 7th pandemic El Tor and related O139 serogroup strains were identified. These latter genes may encode gain-of-function traits specifically associated with displacement of the preexisting classical strains in South Asia and may also promote the establishment of endemic disease in previously cholera-free locations.

Cholera↗

Evolutionary role of restriction/modification systems as revealed by comparative genome analysis.

Type II restriction modification systems (RMSs) have been regarded either as defense tools or as molecular parasites of bacteria. We extensively analyzed their evolutionary role from the study of their impact in the complete genomes of 26 bacteria and 35 phages in terms of palindrome avoidance. This analysis reveals that palindrome avoidance is not universally spread among bacterial species and that it does not correlate with taxonomic proximity. Palindrome avoidance is also not universal among bacteriophage, even when their hosts code for RMSs, and depends strongly on the genetic material of the phage. Interestingly, palindrome avoidance is intimately correlated with the infective behavior of the phage. We observe that the degree of palindrome and restriction site avoidance is significantly and consistently less important in phages than in their bacterial hosts. This result brings to the fore a larger selective load for palindrome and restriction site avoidance on the bacterial hosts than on their infecting phages. It is then consistent with a view where type II RMSs are considered as parasites possibly at the verge of mutualism. As a consequence, RMSs constitute a nontrivial third player in the host-parasite relationship between bacteria and phages.

AT Rich Sequence↗

LegumeDB1 bioinformatics resource: comparative genomic analysis and novel cross-genera marker identification in lupin and pasture legume species.

The identification of markers in legume pasture crops, which can be associated with traits such as protein and lipid production, disease resistance, and reduced pod shattering, is generally accepted as an important strategy for improving the agronomic performance of these crops. It has been demonstrated that many quantitative trait loci (QTLs) identified in one species can be found in other plant species. Detailed legume comparative genomic analyses can characterize the genome organization between model legume species (e.g., Medicago truncatula, Lotus japonicus) and economically important crops such as soybean (Glycine max), pea (Pisum sativum), chickpea (Cicer arietinum), and lupin (Lupinus angustifolius), thereby identifying candidate gene markers that can be used to track QTLs in lupin and pasture legume breeding. LegumeDB is a Web-based bioinformatics resource for legume researchers. LegumeDB analysis of Medicago truncatula expressed sequence tags (ESTs) has identified novel simple sequence repeat (SSR) markers (16 tested), some of which have been putatively linked to symbiosome membrane proteins in root nodules and cell-wall proteins important in plant-pathogen defence mechanisms. These novel markers by preliminary PCR assays have been detected in Medicago truncatula and detected in at least one other legume species, Lotus japonicus, Glycine max, Cicer arietinum, and (or) Lupinus angustifolius (15/16 tested). Ongoing research has validated some of these markers to map them in a range of legume species that can then be used to compile composite genetic and physical maps. In this paper, we outline the features and capabilities of LegumeDB as an interactive application that provides legume genetic and physical comparative maps, and the efficient feature identification and annotation of the vast tracks of model legume sequences for convenient data integration and visualization. LegumeDB has been used to identify potential novel cross-genera polymorphic legume markers that map to agronomic traits, supporting the accelerated identification of molecular genetic factors underpinning important agronomic attributes in lupin.

Chromosome Mapping↗

Comparative genome analysis delimits a chromosomal domain and identifies key regulatory elements in the alpha globin cluster.

We have cloned, sequenced and annotated segments of DNA spanning the mouse, chicken and pufferfish alpha globin gene clusters and compared them with the corresponding region in man. This has defined a small segment ( approximately 135-155 kb) of synteny and conserved gene order, which may contain all of the elements required to fully regulate alpha globin gene expression from its natural chromosomal environment. Comparing human and mouse sequences using previously described methods failed to identify the known regulatory elements. However, refining these methods by ranking identity scores of non-coding sequences, we found conserved sequences including the previously characterized alpha globin major regulatory element. In chicken and pufferfish, regions that may correspond to this element were found by analysing the distribution of transcription factor binding sites. Regions identified in this way act as strong enhancer elements in expression assays. In addition to delimiting the alpha globin chromosomal domain, this study has enabled us to develop a more sensitive and accurate routine for identifying regulatory elements in the human genome.

Animals↗

Comparative genomic analysis of equilibrative nucleoside transporters suggests conserved protein structure despite limited sequence identity.

Equilibrative nucleoside transporters (ENTs) are a recently characterized and poorly understood group of membrane proteins that are important in the uptake of endogenous nucleosides required for nucleic acid and nucleoside triphosphate synthesis. Despite their central importance in cellular metabolism and nucleoside analog chemotherapy, no human ENT gene has been described and nothing is known about gene structure and function. To gain insight into the ENT gene family, we used experimental and in silico comparative genomic approaches to identify ENT genes in three evolutionarily diverse organisms with completely (or almost completely) sequenced genomes, Homo sapiens, Caenorhabditis elegans and Drosophila melanogaster. We describe the chromosomal location, the predicted ENT gene structure and putative structural topologies of predicted ENT proteins derived from the open reading frames. Despite variations in genomic layout and limited ortholog protein sequence identity (< or =27.45%), predicted topologies of ENT proteins are strikingly similar, suggesting an evolutionary conservation of a prototypic structure. In addition, a similar distribution of protein domains on exons is apparent in all three taxa. These data demonstrate that comparative sequence analyses should be combined with other approaches (such as genomic and proteomic analyses) to fully understand structure, function and evolution of protein families.

Alternative Splicing↗

Comparative genomic analysis of Chlamydia trachomatis oculotropic and genitotropic strains.

Chlamydia trachomatis infection is an important cause of preventable blindness and sexually transmitted disease (STD) in humans. C. trachomatis exists as multiple serovariants that exhibit distinct organotropism for the eye or urogenital tract. We previously reported tissue-tropic correlations with the presence or absence of a functional tryptophan synthase and a putative GTPase-inactivating domain of the chlamydial toxin gene. This suggested that these genes may be the primary factors responsible for chlamydial disease organotropism. To test this hypothesis, the genome of an oculotropic trachoma isolate (A/HAR-13) was sequenced and compared to the genome of a genitotropic (D/UW-3) isolate. Remarkably, the genomes share 99.6% identity, supporting the conclusion that a functional tryptophan synthase enzyme and toxin might be the principal virulence factors underlying disease organotropism. Tarp (translocated actin-recruiting phosphoprotein) was identified to have variable numbers of repeat units within the N and C portions of the protein. A correlation exists between lymphogranuloma venereum serovars and the number of N-terminal repeats. Single-nucleotide polymorphism (SNP) analysis between the two genomes highlighted the minimal genetic variation. A disproportionate number of SNPs were observed within some members of the polymorphic membrane protein (pmp) autotransporter gene family that corresponded to predicted T-cell epitopes that bind HLA class I and II alleles. These results implicate Pmps as novel immune targets, which could advance future chlamydial vaccine strategies. Lastly, a novel target for PCR diagnostics was discovered that can discriminate between ocular and genital strains. This discovery will enhance epidemiological investigations in nations where both trachoma and chlamydial STD are endemic.

Amino Acid Sequence↗

Comparative genomic analysis of sequences sampled from a small region on soybean (Glycine max) molecular linkage group G.

Eight DNA markers spanning an interval of approximately 10 centimorgans (cM) on soybean (Glycine max) molecular linkage group G (MLG-G) were used to identify bacterial artificial chromosome (BAC) clones. Twenty-eight BAC clones in eight distinct contiguous groups (contigs) were isolated from this genome region, along with 59 BAC clones on 17 contigs homoeologous to those on MLG-G. BAC clones in four of the MLG-G contigs were also digested to produce subclones and detailed physical maps. All of the BAC-ends were sequenced, as were the subclones, to estimate proportions in different sequence categories, compare similarities among homoeologs, and explore microsynteny with Arabidopsis. Homoeologous BAC contigs were enriched in repetitive sequences compared with those on MLG-G or the soybean genome as a whole. Fingerprint and cross-hybridization comparisons between MLG-G and homoeologous contigs revealed cases of highly similar physical organization between soybean duplicates, as did DNA sequence comparisons. Twenty-seven out of 78 total sequences on soybean MLG-G showed significant similarity to Arabidopsis. The homologs mapped to six compact genome segments in Arabidopsis, with the longest containing seven homologs spanning two million base pairs. These results extend previous observations of large-scale duplication and selective gene loss in Arabidopsis, suggesting that networks of conserved synteny between Arabidopsis and other angiosperm families can stretch over long physical distances.

Arabidopsis↗

Comparative genome analysis of Bacillus cereus group genomes with Bacillus subtilis.

Genome features of the Bacillus cereus group genomes (representative strains of Bacillus cereus, Bacillus anthracis and Bacillus thuringiensis sub spp. israelensis) were analyzed and compared with the Bacillus subtilis genome. A core set of 1381 protein families among the four Bacillus genomes, with an additional set of 933 families common to the B. cereus group, was identified. Differences in signal transduction pathways, membrane transporters, cell surface structures, cell wall, and S-layer proteins suggesting differences in their phenotype were identified. The B. cereus group has signal transduction systems including a tyrosine kinase related to two-component system histidine kinases from B. subtilis. A model for regulation of the stress responsive sigma factor sigmaB in the B. cereus group different from the well studied regulation in B. subtilis has been proposed. Despite a high degree of chromosomal synteny among these genomes, significant differences in cell wall and spore coat proteins that contribute to the survival and adaptation in specific hosts has been identified.

Bacillus anthracis↗