Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comparative genomic analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

How to become a uropathogen: comparative genomic analysis of extraintestinal pathogenic Escherichia coli strains.

Uropathogenic Escherichia coli (UPEC) strain 536 (O6:K15:H31) is one of the model organisms of extraintestinal pathogenic E. coli (ExPEC). To analyze this strain's genetic basis of urovirulence, we sequenced the entire genome and compared the data with the genome sequence of UPEC strain CFT073 (O6:K2:H1) and to the available genomes of nonpathogenic E. coli strain MG1655 (K-12) and enterohemorrhagic E. coli. The genome of strain 536 is approximately 292 kb smaller than that of strain CFT073. Genomic differences between both UPEC are mainly restricted to large pathogenicity islands, parts of which are unique to strain 536 or CFT073. Genome comparison underlines that repeated insertions and deletions in certain parts of the genome contribute to genome evolution. Furthermore, 427 and 432 genes are only present in strain 536 or in both UPEC, respectively. The majority of the latter genes is encoded within smaller horizontally acquired DNA regions scattered all over the genome. Several of these genes are involved in increasing the pathogens' fitness and adaptability. Analysis of virulence-associated traits expressed in the two UPEC O6 strains, together with genome comparison, demonstrate the marked genetic and phenotypic variability among UPEC. The ability to accumulate and express a variety of virulence-associated genes distinguishes ExPEC from many commensals and forms the basis for the individual virulence potential of ExPEC. Accordingly, instead of a common virulence mechanism, different ways exist among ExPEC to cause disease.

Biological Evolution↗

Complete genome sequence and comparative genomic analysis of an emerging human pathogen, serotype V Streptococcus agalactiae.

The 2,160,267 bp genome sequence of Streptococcus agalactiae, the leading cause of bacterial sepsis, pneumonia, and meningitis in neonates in the U.S. and Europe, is predicted to encode 2,175 genes. Genome comparisons among S. agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, and the other completely sequenced genomes identified genes specific to the streptococci and to S. agalactiae. These in silico analyses, combined with comparative genome hybridization experiments between the sequenced serotype V strain 2603 V/R and 19 S. agalactiae strains from several serotypes using whole-genome microarrays, revealed the genetic heterogeneity among S. agalactiae strains, even of the same serotype, and provided insights into the evolution of virulence mechanisms.

Amino Acid Sequence↗

Comparative genomic analysis of the pPT23A plasmid family of Pseudomonas syringae.

Members of the pPT23A plasmid family of Pseudomonas syringae play an important role in the interaction of this bacterial pathogen with host plants. Complete sequence analysis of several pPT23A family plasmids (PFPs) has provided a glimpse of the gene content and virulence function of these plasmids. We constructed a macroarray containing 161 genes to estimate and compare the gene contents of 23 newly analyzed and eight known PFPs from 12 pathovars of P. syringae, which belong to four genomospecies. Hybridization results revealed that PFPs could be distinguished by the type IV secretion system (T4SS) encoded and separated into four groups. Twelve PFPs along with pPSR1 from P. syringae pv. syringae, pPh1448B from P. syringae pv. phaseolicola, and pPMA4326A from P. syringae pv. maculicola encoded a type IVA T4SS (VirB-VirD4 conjugative system), whereas 10 PFPs along with pDC3000A and pDC3000B from P. syringae pv. tomato encoded a type IVB T4SS (tra system). Two plasmids encoded both T4SSs, whereas six other plasmids carried none or only a few genes of either the type IVA or type IVB secretion system. Most PFPs hybridized to more than one putative type III secretion system effector gene and to a variety of additional genes encoding known P. syringae virulence factors. The overall gene contents of individual PFPs were more similar among plasmids within each of the four groups based on T4SS genes; however, a number of genes, encoding plasmid-specific functions or hypothetical proteins, were shared among plasmids from different T4SS groups. The only gene shared by all PFPs in this study was the repA gene, which encoded sequences with 87 to 99% amino acid identityamong 25 sequences examined. We proposed a model to illustrate the evolution and gene acquisition of the pPT23A plasmid family. To our knowledge, this is the first such attempt to conduct a global genetic analysis of this important plasmid family.

DNA Helicases↗

Comparative genomic analysis reveals a distant liver enhancer upstream of the COUP-TFII gene.

COUP-TFII is a central nuclear hormone receptor that tightly regulates the expression of numerous target lipid metabolism genes in vertebrates. However, it remains unclear how COUP-TFII itself is transcriptionally controlled since studies with its promoter and upstream region fail to recapitulate the gene's liver expression. In an attempt to identify liver enhancers in the vicinity of COUP-TFII, we employed a comparative genomic approach. Initial comparisons between humans and mice of the 3470-kb gene-poor region surrounding COUP-TFII revealed 2023 conserved noncoding elements. To prioritize a subset of these elements for functional studies, we performed further genomic comparisons with the orthologous pufferfish (Fugu rubripes) locus and uncovered two anciently conserved noncoding sequences (CNS) upstream of COUP-TFII (CNS-62kb and CNS-66kb). Testing these two elements using reporter constructs in liver cells (HepG2) revealed that CNS-66kb, but not CNS-62kb, yielded robust in vitro enhancer activity. In addition, an in vivo reporter assay using naked DNA transfer with CNS-66kb linked to luciferase displayed strong reproducible liver expression in adult mice, further supporting its role as a liver enhancer. Together, these studies further support the utility of comparative genomics to uncover gene regulatory sequences based on evolutionary conservation and provide the substrates to better understand the regulation and expression of COUP-TFII.

Animals↗

Comparative genomic analysis of dha regulon and related genes for anaerobic glycerol metabolism in bacteria.

The dihydroxyacetone (dha) regulon of bacteria encodes genes for the anaerobic metabolism of glycerol. In this work, genomic data are used to analyze and compare the dha regulon and related genes in different organisms in silico with respect to gene organization, sequence similarity, and possible functions. Database searches showed that among the organisms, the genomes of which have been sequenced so far, only two, i.e., Klebsiella pneumoniae MGH 78578 and Clostridium perfringens contain a complete dha regulon bearing all known enzymes. The components and their organization in the dha regulon of these two organisms differ considerably from each other and also from the previously partially sequenced dha regulons in Citrobacter freundii, Clostridium pasteurianum, and Clostridium butyricum. Unlike all of the other organisms, genes for the oxidative and reductive pathways of anaerobic glycerol metabolism in C. perfringens are located in two separate organization units on the chromosome. Comparisons of deduced protein sequences of genes with similar functions showed that the dha regulon components in K. pneumoniae and C. freundii have high similarities (80-95%) but lower similarities to those of the Clostridium species (30-80%). Interestingly, the protein sequence similarities among the dha genes of the Clostridium species are in many cases even lower than those between the Clostridium species and K. pneumoniae or C. freundii, suggesting two different types of dha regulon in the Clostridium species studied. The in silico reconstruction and comparison of dha regulons revealed several new genes in the microorganisms studied. In particular, a novel dha kinase that is phosphoenolpyruvate-dependent is identified and experimentally confirmed for K. pneumoniae in addition to the known ATP-dependent dha kinase. This finding gives new insights into the regulation of glycerol metabolism in K. pneumoniae and explains some hitherto not well understood experimental observations.

Amino Acid Sequence↗

Comparative genomic analysis of allatostatin-encoding (Ast) genes in Drosophila species and prediction of regulatory elements by phylogenetic footprinting.

The role of the YXFGLa family of allatostatin (AST) peptides in dipterans is not well-established. The recent completion of sequencing of genomes for multiple Drosophila species provides an opportunity to study the evolutionary variation of the allatostatins and to examine regulatory elements that control gene expression. We performed comparative analyses of Ast genes from seven Drosophila species (Drosophila melanogaster, Drosophila simulans, Drosophila ananassae, Drosophila yakuba, Drosophila pseudoobscura, Drosophila mojavensis, and Drosophila grimshawi) and used phylogenetic footprinting methods to identify conserved noncoding motifs, which are candidates for regulatory regions. The peptides encoded by the Ast precursor are nearly identical across species with the exception of AST-1, in which the leading residue may be either methionine or valine. Phylogenetic footprinting predicts as few as 3, to as many as 17 potential regulatory sites depending on the parameters used during analysis. These include a Hunchback motif approximately 1.2 kb upstream of the open reading frame (ORF), overlapping motifs for two Broad-complex isoforms in the first intron, and a CF2-II motif located in the 3'-UTR. Understanding the regulatory elements involved in Ast expression may provide insight into the function of this neuropeptide family.

Amino Acid Sequence↗

A comparative genomic analysis of the cow, pig, and human CFTR genes identifies potential intronic regulatory elements.

The identification of sequences within noncoding regions of genes that are conserved between several species may indicate potential regulatory elements. This is important for genes with complex control mechanisms such as the cystic fibrosis transmembrane conductance regulator (CFTR). CFTR demonstrates similar patterns of temporal and spatial expression in human and sheep, but these differ significantly in mouse cftr. The complete sheep CFTR sequence is unavailable so we annotated BAC clones encompassing the CFTR gene from two other artiodactyl species (cow and pig) for comparative sequence analysis. Regions of introns 2, 3, 10, 17a, 18, and 21 and 3' flanking sequence corresponding to human CFTR DNase I hypersensitive sites (DHS) showed high homology in the cow and pig. Cross-species sequence conservation also enabled finer mapping of other human DHS, including those in introns 1, 16, and 20. Additional potential regulatory elements not associated with human DHS were also identified.

Animals↗

Comparative genomic analysis of human and chimpanzee proteases.

Proteolytic enzymes are implicated in multiple physiological and pathological processes. The availability of the sequence of the chimpanzee genome has allowed us to determine that the chimpanzee degradome-the repertoire of protease genes from this organism-is composed of at least 559 protease and protease-like genes and is virtually identical to that of human, containing 561 genes. Despite the high degree of conservation between both genomes, we have identified important differences that vary from deletion of whole genes to small insertion/deletion events or single nucleotide changes that lead to the specific gene inactivation in one species, mostly affecting immune system genes. For example, the genes encoding PRSS33/EOS, a macrophage serine protease conserved in most mammals, and GGTLA1 are absent in chimpanzee, while the gene for metalloprotease MMP23A, located in chromosome 1p36, has been specifically duplicated in the human genome together with its neighbor gene CDC2L1. Other differences arise from single nucleotide changes in protease genes, such as NAPSB and CASP12, resulting in the presence of functional genes in chimpanzee and pseudogenes in human. Finally, we have confirmed that the Trypanosoma lytic factor HPR is inactive in chimpanzee, likely contributing to the susceptibility of chimpanzees to T. brucei infection. This study provides the first analysis of the chimpanzee degradome and might contribute to the understanding of the molecular bases underlying variations in host defense mechanisms between human and chimpanzee.

Amino Acid Sequence↗

Origin and evolutionary process of the CNS elucidated by comparative genomics analysis of planarian ESTs.

Among the bilateral animals, a centralized nervous system is found in both the deuterostome and protostome. To address the question of whether the CNS was derived from a common ancestor of deuterostomes and protostomes, it is essential to know kinds of genes existed in the CNS of the putative common ancestor and to trace the evolutionary divergence of genes expressed in the CNS. To answer these questions, we took a comparative approach using different species, particularly focusing on one of the lower bilateral animals, the planarian (Platyhelminthes, Tricladida), which is known to possess a CNS. We determined the nucleotide sequence of ESTs from the head portion of planarians, obtaining 3,101 nonredundant EST clones. As a result of homology searches, we found that 116 clones had significant similarity to known genes related to the nervous system. Here, we compared these 116 planarian EST clones with all ORFs of the complete genome sequences of the human, fruit fly, and nematode, and showed that >95% of these 116 nervous system-related genes, including genes involved in brain or neural morphogenesis, were commonly shared among these organisms, thus providing evidence at the molecular level for the existence of a common ancestral CNS. Interestingly, we found that approximately 30% of planarian nervous system-related genes had homologous sequences in Arabidopsis and yeast, which do not possess a nervous system. This implies that the origin of nervous system-related genes greatly predated the emergence of the nervous system, and that these genes might have been recruited toward the nervous system.

Animals↗

Comparative genomic analysis of transcription regulation elements involved in human map kinase G-protein coupling pathway.

The identification of cis-elements (motifs) in the regulatory regions of higher eukaryotes is an important and challenging problem in computational biology. Eukaryotic transcriptional regulatory mechanisms pose several difficulties for promoter analysis: including a high variance in the motif locations, frequently large divergence from motif consensus patterns, and a large amount of repetitive elements (confusing to many motif finding procedures). One promising approach to this difficult problem involves cross-species comparison. In this work we analyzed the full-length regulatory regions of genes involved in the G-protein coupling MAP kinase pathway and compared the results with ribosomal genes using human, mouse and rat genomic data. We found 19 high likely transcription factors (TFs) candidates for MAPK and 12 TFs for the ribosomal dataset. In the case of the MAPK dataset, regulatory regions of genes functionally grouped as receptors and MAP-core genes were found mostly highly conserved across the three species.

Animals↗

Comparative genomic analysis reveals independent expansion of a lineage-specific gene family in vertebrates: the class II cytokine receptors and their ligands in mammals and fish.

BACKGROUND: The high degree of sequence conservation between coding regions in fish and mammals can be exploited to identify genes in mammalian genomes by comparison with the sequence of similar genes in fish. Conversely, experimentally characterized mammalian genes may be used to annotate fish genomes. However, gene families that escape this principle include the rapidly diverging cytokines that regulate the immune system, and their receptors. A classic example is the class II helical cytokines (HCII) including type I, type II and lambda interferons, IL10 related cytokines (IL10, IL19, IL20, IL22, IL24 and IL26) and their receptors (HCRII). Despite the report of a near complete pufferfish (Takifugu rubripes) genome sequence, these genes remain undescribed in fish. RESULTS: We have used an original strategy based both on conserved amino acid sequence and gene structure to identify HCII and HCRII in the genome of another pufferfish, Tetraodon nigroviridis that is amenable to laboratory experiments. The 15 genes that were identified are highly divergent and include a single interferon molecule, three IL10 related cytokines and their potential receptors together with two Tissue Factor (TF). Some of these genes form tandem clusters on the Tetraodon genome. Their expression pattern was determined in different tissues. Most importantly, Tetraodon interferon was identified and we show that the recombinant protein can induce antiviral MX gene expression in Tetraodon primary kidney cells. Similar results were obtained in Zebrafish which has 7 MX genes. CONCLUSION: We propose a scheme for the evolution of HCII and their receptors during the radiation of bony vertebrates and suggest that the diversification that played an important role in the fine-tuning of the ancestral mechanism for host defense against infections probably followed different pathways in amniotes and fish.

Animals↗

[Comparative genome analysis of Helicobacter pylori strains].

DNA macroarrays were used to characterize 17 Helicobacter pylori strains isolated in four geographic regions of Russia (Moscow, St. Petersburg, Kazan, and Novosibirsk). Of all genes, 1272 (81%) proved to occur in all strains and to constitute a functional core of the genome, and 293 (18.7%) were strain-specific and greatly varied among the H. pylori strains. Most (71%) of the latter had unknown functions; the remainder included restriction-modification genes (3-9%), transposition genes (2-4%), and genes coding for outer membrane proteins (2-4%). The Russian H. pylori strains did not differ in genome organization or in the number and distribution of strain-specific genes from strains isolated in other countries.

Bacterial Proteins↗

Complete comparative genomic analysis of two field isolates of Mamestra configurata nucleopolyhedrovirus-A.

A second genotype of Mamestra configurata nucleopolyhedrovirus-A (MacoNPV-A), variant 90/4 (v90/4), was identified due to its altered restriction endonuclease profile and reduced virulence for the host insect, M. configurata, relative to the archetypal genotype, MacoNPV-A variant 90/2 (v90/2). To investigate the genetic differences between these two variants, the genome of v90/4 was sequenced completely. The MacoNPV-A v90/4 genome is 153 656 bp in size, 1404 bp smaller than the v90/2 genome. Sequence alignment showed that there was 99.5 % nucleotide sequence identity between the genomes of v90/4 and v90/2. However, the v90/4 genome has 521 point mutations and numerous deletions and insertions when compared to the genome of v90/2. Gene content and organization in the genome of v90/4 is identical to that in v90/2, except for an additional bro gene that is found in the v90/2 genome. The region between hr1 and orf31 shows the greatest divergence between the two genomes. This region contains three bro genes, which are among the most variable baculovirus genes. These results, together with other published data, suggest that bro genes may influence baculovirus genome diversity and may be involved in recombination between baculovirus genomes. Many ambiguous residues found in the v90/4 sequence also reveal the presence of 214 sequence polymorphisms. Sequence analysis of cloned HindIII fragments of the original MacoNPV field isolate that the 90/4 variant was derived from indicates that v90/4 is an authentic variant and may represent approximately 25 % of the genotypes in the field isolate. These results provide evidence of extensive sequence variation among the individual genomes comprising a natural baculovirus outbreak in a continuous host population.

Animals↗

Comparative genomic analysis of human and chimpanzee indicates a key role for indels in primate evolution.

Sequence comparison of humans and chimpanzees is of interest to understand the mechanisms behind primate evolution. Here we present an independent analysis of human chromosome 21 and the high-quality BAC clone sequences of the homologous chimpanzee chromosome 22. In contrast to previous studies, we have used global alignment methods and Ensembl predictions of protein coding genes (n = 224) for the analysis. Divergence due to insertions and deletions (indels) along with substitutions was examined separately for different genomic features (coding, noncoding genic, and intergenic sequence). The major part of the genomic divergence could be attributed to indels (5.07%), while the nucleotide divergence was estimated as 1.52%. Thus the total divergence was estimated as 6.58%. When excluding repeats and low-complexity DNA the total divergence decreased to 2.37%. The chromosomal distribution of nucleotide substitutions and indel events was significantly correlated. To further examine the role of indels in primate evolution we focused on coding sequences. Indels were found within the coding sequence of 13% of the genes and approximately half of the indels have not been reported previously. In 5% of the chimpanzee genes, indels or substitutions caused premature stop codons that rendered the affected transcripts nonfunctional. Taken together, our findings demonstrate that indels comprise the majority of the genomic divergence. Furthermore, indels occur frequently in coding sequences. Our results thereby support the hypothesis that indels may have a key role in primate evolution.

Animals↗

Large-scale genetic variation of the symbiosis-required megaplasmid pSymA revealed by comparative genomic analysis of Sinorhizobium meliloti natural strains.

BACKGROUND: Sinorhizobium meliloti is a soil bacterium that forms nitrogen-fixing nodules on the roots of leguminous plants such as alfalfa (Medicago sativa). This species occupies different ecological niches, being present as a free-living soil bacterium and as a symbiont of plant root nodules. The genome of the type strain Rm 1021 contains one chromosome and two megaplasmids for a total genome size of 6 Mb. We applied comparative genomic hybridisation (CGH) on an oligonucleotide microarrays to estimate genetic variation at the genomic level in four natural strains, two isolated from Italian agricultural soil and two from desert soil in the Aral Sea region. RESULTS: From 4.6 to 5.7 percent of the genes showed a pattern of hybridisation concordant with deletion, nucleotide divergence or ORF duplication when compared to the type strain Rm 1021. A large number of these polymorphisms were confirmed by sequencing and Southern blot. A statistically significant fraction of these variable genes was found on the pSymA megaplasmid and grouped in clusters. These variable genes were found to be mainly transposases or genes with unknown function. CONCLUSION: The obtained results allow to conclude that the symbiosis-required megaplasmid pSymA can be considered the major hot-spot for intra-specific differentiation in S. meliloti.

Blotting, Southern↗

Comparative genome analysis of the neurexin gene family in Danio rerio: insights into their functions and evolution.

Neurexins constitute a family of proteins originally identified as synaptic transmembrane receptors for a spider venom toxin. In mammals, the 3 known Neurexin genes present 2 alternative promoters that drive the synthesis of a long (alpha) and a short (beta) form and contain different sites of alternative splicing (AS) that can give rise to thousands of different transcripts. To date, very little is known about the significance of this variability, except for the modulation of binding to some of the Neurexin ligands. Although orthologs of Neurexins have been isolated in invertebrates, these genes have been studied mostly in mammals. With the aim of investigating their functions in lower vertebrates, we chose Danio rerio as a model because of its increasing importance in comparative biology. We have isolated 6 zebrafish homologous genes, which are highly conserved at the structural level and display a similar regulation of AS, despite about 450 Myr separating the human and zebrafish species. Our data indicate a strong selective pressure at the exonic level and on the intronic borders, in particular on the regulative intronic sequences that flank the exons subject to AS. Such a selective pressure could help conserve the regulation and consequently the function of these genes along the vertebrates evolutive tree. AS analysis during development shows that all genes are expressed and finely regulated since the earliest stages of development, but mark an increase after the 24-h stage that corresponds to the beginning of synaptogenesis. Moreover, we found that specific isoforms of a zebrafish Neurexin gene (nrxn1a) are expressed in the adult testis and in the earliest stages of development, before the beginning of zygotic transcription, indicating a potential delivery of paternal RNA to the embryo. Our analysis suggests the existence of possible new functions for Neurexins, serving as the basis for novel approaches to the functional studies of this complex neuronal protein family and more in general to the understanding of the AS mechanism in low vertebrates.

Alternative Splicing↗