Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comparative genomic analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Comparative genomic analysis of three strains of Ehrlichia ruminantium reveals an active process of genome size plasticity.

Ehrlichia ruminantium is the causative agent of heartwater, a major tick-borne disease of livestock in Africa that has been introduced in the Caribbean and is threatening to emerge and spread on the American mainland. We sequenced the complete genomes of two strains of E. ruminantium of differing phenotypes, strains Gardel (Erga; 1,499,920 bp), from the island of Guadeloupe, and Welgevonden (Erwe; 1,512,977 bp), originating in South Africa and maintained in Guadeloupe in a different cell environment. Comparative genomic analysis of these two strains was performed with the recently published parent strain of Erwe (Erwo) and other Rickettsiales (Anaplasma, Wolbachia, and Rickettsia spp.). Gene order is highly conserved between the E. ruminantium strains and with A. marginale. In contrast, there is very little conservation of gene order with members of the Rickettsiaceae. However, gene order may be locally conserved, as illustrated by the tuf operons. Eighteen truncated protein-encoding sequences (CDSs) differentiate Erga from Erwe/Erwo, whereas four other truncated CDSs differentiate Erwe from Erwo. Moreover, E. ruminantium displays the lowest coding ratio observed among bacteria due to unusually long intergenic regions. This is related to an active process of genome expansion/contraction targeted at tandem repeats in noncoding regions and based on the addition or removal of ca. 150-bp tandem units. This process seems to be specific to E. ruminantium and is not observed in the other Rickettsiales.

Conserved Sequence↗

Comparative genomic analysis revealed a gene for monoglucosyldiacylglycerol synthase, an enzyme for photosynthetic membrane lipid synthesis in cyanobacteria.

Cyanobacteria have a thylakoid lipid composition very similar to that of plant chloroplasts, yet cyanobacteria are proposed to synthesize monogalactosyldiacylglycerol (MGDG), a major membrane polar lipid in photosynthetic membranes, by a different pathway. In addition, plant MGDG synthase has been cloned, but no ortholog has been reported in cyanobacterial genomes. We report here identification of the gene for monoglucosyldiacylglycerol (MGlcDG) synthase, which catalyzes the first step of galactolipid synthesis in cyanobacteria. Using comparative genomic analysis, candidates for the gene were selected based on the criteria that the enzyme activity is conserved between two species of cyanobacteria (unicellular [Synechocystis sp. PCC 6803] and filamentous [Anabaena sp. PCC 7120]), and we assumed three characteristics of the enzyme; namely, it harbors a glycosyltransferase motif, falls into a category of genes with unknown function, and shares significant similarity in amino acid sequence between these two cyanobacteria. By a motif search of all genes of Synechocystis, BLAST searches, and similarity searches between these two cyanobacteria, we identified four candidates for the enzyme that have all the characteristics we predicted. When expressed in Escherichia coli, one of the Synechocystis candidate proteins showed MGlcDG synthase activity in a UDP-glucose-dependent manner. The ortholog in Anabaena also showed the same activity. The enzyme was predicted to require a divalent cation for its activity, and this was confirmed by biochemical analysis. The MGlcDG synthase and the plant MGDG synthase shared low similarity, supporting the presumption that cyanobacteria and plants utilize different pathways to synthesize MGDG.

Amino Acid Sequence↗

Identification of cyanobacterial non-coding RNAs by comparative genome analysis.

BACKGROUND: Whole genome sequencing of marine cyanobacteria has revealed an unprecedented degree of genomic variation and streamlining. With a size of 1.66 megabase-pairs, Prochlorococcus sp. MED4 has the most compact of these genomes and it is enigmatic how the few identified regulatory proteins efficiently sustain the lifestyle of an ecologically successful marine microorganism. Small non-coding RNAs (ncRNAs) control a plethora of processes in eukaryotes as well as in bacteria; however, systematic searches for ncRNAs are still lacking for most eubacterial phyla outside the enterobacteria. RESULTS: Based on a computational prediction we show the presence of several ncRNAs (cyanobacterial functional RNA or Yfr) in several different cyanobacteria of the Prochlorococcus-Synechococcus lineage. Some ncRNA genes are present only in two or three of the four strains investigated, whereas the RNAs Yfr2 through Yfr5 are structurally highly related and are encoded by a rapidly evolving gene family as their genes exist in different copy numbers and at different sites in the four investigated genomes. One ncRNA, Yfr7, is present in at least seven other cyanobacteria. In addition, control elements for several ribosomal operons were predicted as well as riboswitches for thiamine pyrophosphate and cobalamin. CONCLUSION: This is the first genome-wide and systematic screen for ncRNAs in cyanobacteria. Several ncRNAs were both computationally predicted and their presence was biochemically verified. These RNAs may have regulatory functions and each shows a distinct phylogenetic distribution. Our approach can be applied to any group of microorganisms for which more than one total genome sequence is available for comparative analysis.

Base Sequence↗

Utility of comparative anchor-tagged sequences as physical anchors for comparative genome analysis among the Culicidae.

The development of comparative genetic maps in multiple species of mosquitoes could prove extremely useful in the search for those genes that contribute to mosquito vector competence or genes associated with other phenotypes of interest. To effectively compare these gene maps, markers must be developed that are based on chromosomal regions conserved throughout the Culicidae. We designed 35 polymerase chain reaction (PCR) primer pairs based upon orthlogous exons in Aedes aegypti and Drosophila melanogaster or Anopheles gambiae. Twenty-three of the primers yielded a single PCR product in at least one dipteran, in addition to Ae. aegypti, when screened with genomic DNA from seven dipterans, including five mosquito species. Eight of the primers amplified a single PCR product in only Ae. aegypti, while four primer pairs gave no PCR product in any species. The 23 successful comparative anchor-tagged sequence primer pairs give broad genome coverage in Ae. aegypti, and more importantly demonstrate an efficient strategy for developing comparative anchor marker loci for any species of Culicidae.

Amino Acid Sequence↗

A comparative genome analysis identifies distinct sorting pathways in gram-positive bacteria.

Surface proteins in gram-positive bacteria are frequently required for virulence, and many are attached to the cell wall by sortase enzymes. Bacteria frequently encode more than one sortase enzyme and an even larger number of potential sortase substrates that possess an LPXTG-type cell wall sorting signal. In order to elucidate the sorting pathways present in gram-positive bacteria, we performed a comparative analysis of 72 sequenced microbial genomes. We show that sortase enzymes can be partitioned into five distinct subfamilies based upon their primary sequences and that most of their substrates can be predicted by making a few conservative assumptions. Most bacteria encode sortases from two or more subfamilies, which are predicted to function nonredundantly in sorting proteins to the cell surface. Only approximately 20% of sortase-related proteins are most closely related to the well-characterized Staphylococcus aureus SrtA protein, but nonetheless, these proteins are responsible for anchoring the majority of surface proteins in gram-positive bacteria. In contrast, most sortase-like proteins are predicted to play a more specialized role, with each anchoring far fewer proteins that contain unusual sequence motifs. The functional sortase-substrate linkage predictions are available online (http://www.doe-mbi.ucla.edu/Services/Sortase/) in a searchable database.

Amino Acid Motifs↗

The complete genome sequence and comparative genome analysis of the high pathogenicity Yersinia enterocolitica strain 8081.

The human enteropathogen, Yersinia enterocolitica, is a significant link in the range of Yersinia pathologies extending from mild gastroenteritis to bubonic plague. Comparison at the genomic level is a key step in our understanding of the genetic basis for this pathogenicity spectrum. Here we report the genome of Y. enterocolitica strain 8081 (serotype 0:8; biotype 1B) and extensive microarray data relating to the genetic diversity of the Y. enterocolitica species. Our analysis reveals that the genome of Y. enterocolitica strain 8081 is a patchwork of horizontally acquired genetic loci, including a plasticity zone of 199 kb containing an extraordinarily high density of virulence genes. Microarray analysis has provided insights into species-specific Y. enterocolitica gene functions and the intraspecies differences between the high, low, and nonpathogenic Y. enterocolitica biotypes. Through comparative genome sequence analysis we provide new information on the evolution of the Yersinia. We identify numerous loci that represent ancestral clusters of genes potentially important in enteric survival and pathogenesis, which have been lost or are in the process of being lost, in the other sequenced Yersinia lineages. Our analysis also highlights large metabolic operons in Y. enterocolitica that are absent in the related enteropathogen, Yersinia pseudotuberculosis, indicating major differences in niche and nutrients used within the mammalian gut. These include clusters directing, the production of hydrogenases, tetrathionate respiration, cobalamin synthesis, and propanediol utilisation. Along with ancestral gene clusters, the genome of Y. enterocolitica has revealed species-specific and enteropathogen-specific loci. This has provided important insights into the pathology of this bacterium and, more broadly, into the evolution of the genus. Moreover, wider investigations looking at the patterns of gene loss and gain in the Yersinia have highlighted common themes in the genome evolution of other human enteropathogens.

Evolution, Molecular↗

Comparative genomic analysis of the HNF-4alpha transcription factor gene.

Hepatocyte nuclear factor-4alpha (HNF-4alpha), the gene for the maturity-onset diabetes of the young type 1 (MODY1) form of type 2 diabetes mellitus (T2DM), is within the T2DM-linked region on chromosome 20q12-q13.1 and consequently, is a positional candidate gene for T2DM. Mutations in the coding region of HNF-4alpha are rare in diabetes affected subjects. Altered regulation of HNF-4alpha gene expression, controlled by distant enhancer sequences, may contribute to the development of type 2 diabetes. Comparative sequence analysis was performed between 13 kb of genomic DNA 5' to the P1 promoter sequences of the human, mouse, and rat HNF-4alpha coding sequences. Three regions, located at -10.5 kb (295 bp in length), -6.25 kb (421 bp in length), and -5.36 kb (263 bp in length), have significant sequence identity between the species. These three regions were functionally characterized using the chloramphenicol acetyltransferase (CAT) reporter assay, in which the conserved 5' regions of mouse HNF-4alpha were cloned in front of the herpes simplex virus thymidine kinase promoter driving transcription of the CAT gene. A fragment containing the 421 bp conserved region significantly increased CAT activity in differentiated rat hepatoma cells (13.7-+/-1.9-fold control), while only a modest increase in CAT activity was observed in pancreatic cells (2.5-+/-0.9-fold control; 1.6-+/-0.1-fold control) and dedifferentiated hepatoma cells (1.7-+/-0.4-fold control). The remaining two conserved regions increased CAT activity minimally in pancreatic (1.1-+/-0.1-fold control to 1.9-+/-0.1-fold control) and hepatic (1.6-+/-0.5-fold control to 2.3-+/-0.4-fold control) cell lines. Denaturing high-performance liquid chromatography (DHPLC) was used to search for sequence variants in DNA from 259 T2DM individuals. Two single nucleotide polymorphisms (SNPs) were identified, both of which increased CAT activity in the insulinoma cell lines in the CAT reporter assay (1.4-fold increase over wild-type; 1.7-fold increase over wild-type). These results suggest that comparative sequence analysis can efficiently identify regulatory elements and that sequence variants in regulatory elements of HNF-4alpha can contribute to altered HNF-4alpha gene expression.

Amino Acid Sequence↗

Comparative genome analysis of the primary sex-determining locus in salmonid fishes.

We compared the Y-chromosome linkage maps for four salmonid species (Arctic charr, Salvelinus alpinus; Atlantic salmon, Salmo salar; brown trout, Salmo trutta; and rainbow trout, Oncorhynchus mykiss) and a putative Y-linked marker from lake trout (Salvelinus namaycush). These species represent the three major genera within the subfamily Salmoninae of the Salmonidae. The data clearly demonstrate that different Y-chromosomes have evolved in each of the species. Arrangements of markers proximal to the sex-determining locus are preserved on homologous, but different, autosomal linkage groups across the four species studied in detail. This indicates that a small region of DNA has been involved in the rearrangement of the sex-determining region. Placement of the sex-determining region appears telomeric in brown trout, Atlantic salmon, and Arctic charr, whereas an intercalary location for SEX may exist in rainbow trout. Three hypotheses are proposed to account for the relocation: translocation of a small chromosome arm; transposition of the sex-determining gene; or differential activation of a primary sex-determining gene region among the species.

Animals↗

Comparative genome analysis of the yellow fever mosquito Aedes aegypti with Drosophila melanogaster and the malaria vector mosquito Anopheles gambiae.

An in silico comparative genomics approach was used to identify putative orthologs to genetically mapped genes from the mosquito, Aedes aegypti, in the Drosophila melanogaster and Anopheles gambiae genome databases. Comparative chromosome positions of 73 D. melanogaster orthologs indicated significant deviations from a random distribution across each of the five A. aegypti chromosomal regions, suggesting that some ancestral chromosome elements have been conserved. However, the two genomes also reflect extensive reshuffling within and between chromosomal regions. Comparative chromosome positions of A. gambiae orthologs indicate unequivocally that A. aegypti chromosome regions share extensive homology to the five A. gambiae chromosome arms. Whole-arm or near-whole-arm homology was contradicted with only two genes among the 75 A. aegypti genes for which orthologs to A. gambiae were identified. The two genomes contain large conserved chromosome segments that generally correspond to break/fusion events and a reciprocal translocation with extensive paracentric inversions evident within. Only very tightly linked genes are likely to retain conserved linear orders within chromosome segments. The D. melanogaster and A. gambiae genome databases therefore offer limited potential for comparative positional gene determinations among even closely related dipterans, indicating the necessity for additional genome sequencing projects with other dipteran species.

Aedes↗

Comparative genome analysis in American marsupials: chromosome banding and in-situ hybridization.

We performed a comparative analysis of the G- and C-banded karyotypes of seven species of didelphid marsupials, representing the three diploid numbers (2n = 14, 18 and 22) known to occur in this family. In addition to a great similarity among karyotypes with the same diploid numbers, we also identified homeologies for all autosomal arms comprising the three karyotypes. Robertsonian rearrangements, pericentric inversions and heterochromatin variation account for the differences among the karyotypes. Interspecific variation in the size of the sex chromosomes is due to differences in heterochromatic content. In-situ hybridization with total genomic DNA revealed considerable conservation of the euchromatic portions of the three karyotypes and indicated divergence of repetitive DNA sequences in autosomal heterochromatin.

Animals↗

Comparative genomic analysis links karyotypic evolution with genomic evolution in the Indian muntjac (Muntiacus muntjak vaginalis).

The karyotype of Indian muntjacs (Muntiacus muntjak vaginalis) has been greatly shaped by chromosomal fusion, which leads to its lowest diploid number among the extant known mammals. We present, here, comparative results based on draft sequences of 37 bacterial artificial clones (BAC) clones selected by chromosome painting for this special muntjac species. Sequence comparison on these BAC clones uncovered sequence syntenic relationships between the muntjac genome and those of other mammals. We found that the muntjac genome has peculiar features with respect to intron size and evolutionary rates of genes. Inspection of more than 80 pairs of orthologous introns from 15 genes reveals a significant reduction in intron size in the Indian muntjac compared to that of human, mouse, and dog. Evolutionary analysis using 19 genes indicates that the muntjac genes have evolved rapidly compared to other mammals. In addition, we identified and characterized sequence composition of the first BAC clone containing a chromosomal fusion site. Our results shed new light on the genome architecture of the Indian muntjac and suggest that chromosomal rearrangements have been accompanied by other salient genomic changes.

Animals↗

Comparative genomic analysis of two strains of human adenovirus type 3 isolated from children with acute respiratory infection in southern China.

Human adenovirus type 3 (HAdV-3) is a causative agent of acute respiratory disease, which is prevalent throughout the world, especially in Asia. Here, the complete genome sequences of two field strains of HAdV-3 (strains GZ1 and GZ2) isolated from children with acute respiratory infection in southern China are reported (GenBank accession nos DQ099432 and DQ105654, respectively). The genomes were 35,273 bp (GZ1) and 35,269 bp (GZ2) and both had a G+C content of 51 mol%. They shared 99% nucleotide identity and the four early and five late regions that are characteristic of human adenoviruses. Thirty-nine protein- and two RNA-coding sequences were identified in the genome sequences of both strains. Protein pX had a predicted molecular mass of 8.3 kDa in strain GZ1; this was lower (7.6 kDa) in strain GZ2. Both strains contained 10 short inverted repeats, in addition to their inverted terminal repeats (111 bp). Comparative whole-genome analysis revealed 93 mismatches and four insertions/deletions between the two strains. Strain GZ1 infection produced a typical cytopathic effect, whereas strain GZ2 did not; non-synonymous substitutions in proteins of GZ2 may be responsible for this difference.

Acute Disease↗

Comparative genome analysis of Campylobacter jejuni using whole genome DNA microarrays.

Whole genome DNA microarrays were constructed and used to investigate genomic diversity in 18 Campylobacter jejuni strains from diverse sources. New algorithms were developed that dynamically determine the boundary between the conserved and variable genes. Seven hypervariable plasticity regions (PR) were identified in the genome (PR1 to PR7) containing 136 genes (50%) of the variable gene pool. When comparisons were made with the sequenced strain NCTC11168, the number of absent or divergent genes ranged from 2.6% (40 genes) to 10.2% (163) and in total 16.3% (269) of the genes were variable. PR1 contains genes important in the utilisation of alternative electron acceptors for respiration and may confer a selective advantage to strains in restricted oxygen environments. PR2, 3 and 7 contain many outer membrane and periplasmic proteins and hypothetical proteins of unknown function that might be linked to phenotypic variation and adaptation to different ecological niches. PR4, 5 and 6 contain genes involved in the production and modification of antigenic surface structures.

Algorithms↗

Comparative genomic analysis of 18 Pseudomonas aeruginosa bacteriophages.

A genomic analysis of 18 P. aeruginosa phages, including nine newly sequenced DNA genomes, indicates a tremendous reservoir of proteome diversity, with 55% of open reading frames (ORFs) being novel. Comparative sequence analysis and ORF map organization revealed that most of the phages analyzed displayed little relationship to each other.

DNA, Viral↗

Comparative genomic analysis of archaeal genotypic variants in a single population and in two different oceanic provinces.

Planktonic crenarchaeotes are present in high abundance in Antarctic winter surface waters, and they also make up a large proportion of total cell numbers throughout deep ocean waters. To better characterize these uncultivated marine crenarchaeotes, we analyzed large genome fragments from individuals recovered from a single Antarctic picoplankton population and compared them to those from a representative obtained from deeper waters of the temperate North Pacific. Sequencing and analysis of the entire DNA insert from one Antarctic marine archaeon (fosmid 74A4) revealed differences in genome structure and content between Antarctic surface water and temperate deepwater archaea. Analysis of the predicted gene products encoded by the 74A4 sequence and those derived from a temperate, deepwater planktonic crenarchaeote (fosmid 4B7) revealed many typical archaeal proteins but also several proteins that so far have not been detected in archaea. The unique fraction of marine archaeal genes included, among others, those for a predicted RNA-binding protein of the bacterial cold shock family and a eukaryote-type Zn finger protein. Comparison of closely related archaea originating from a single population revealed significant genomic divergence that was not evident from 16S rRNA sequence variation. The data suggest that considerable functional diversity may exist within single populations of coexisting microbial strains, even those with identical 16S rRNA sequences. Our results also demonstrate that genomic approaches can provide high-resolution information relevant to microbial population genetics, ecology, and evolution, even for microbes that have not yet been cultivated.

Amino Acid Sequence↗

Comparative genomic analysis of the clade B serpin cluster at human chromosome 18q21: amplification within the mouse squamous cell carcinoma antigen gene locus.

The human clade B serpins neutralize serine or cysteine proteinases and reside predominantly within the intracellular compartment. Genomic analysis shows that the 13 human clade B serpins map to either 6p25 (n = 3) or 18q21 (n = 10). Similarly, the mouse clade B serpins map to syntenic loci at 13A3.2 and 1D, respectively. The mouse clade B cluster at 13A3.2 shows a marked expansion in the number of serpin genes (n = 15). The purpose of this study was to determine whether a similar expansion occurred at 1D. Using STS-content mapping, comparative genomic DNA sequence analysis, and cDNA cloning, we found that the mouse clade B cluster at 1D showed nearly complete conservation of gene number, order, and orientation relative to those of 18q21. The only exception was the squamous cell carcinoma antigen (SCCA) locus. The human SCCA locus contains two genes, SERPINB3 (SCCA1) and SERPINB4 (SCCA2), whereas the mouse locus contains four serpins and three pseudogenes. Based on phylogenetic analysis and predicted amino acid sequences, amplification of the mouse SCCA locus occurred after rodents and primates diverged and was associated with some diversification of proteinase inhibitory activity relative to that of humans.

Amino Acid Sequence↗

Comparative Genomic Analysis of Multidrug-Resistant Escherichia coli Across Poultry-Human-Environmental Interfaces.

The emergence of multidrug-resistant (MDR) Escherichia coli in poultry represents a critical One Health concern, particularly in developing countries. This study employed a comparative genomic approach to investigate the genomic characteristics, antimicrobial resistance (AMR) profiles, virulence determinants, of poultry-derived MDR E. coli isolates from Bangladesh. Whole-genome sequencing of three representative MDR isolates, identified with 83 globally diverse poultry, human, and environmental E. coli genomes. Pangenome analysis identified the characteristic open pangenome of E. coli, with core genes comprising only 4.6% of the combined dataset. Resistome analysis shown diverse AMR determinants, including blaCTX-M, blaTEM, sul, tet, and qnrS1, associated with antibiotic inactivation and efflux mechanisms. Virulence profiling revealed diverse genes involved in adhesion (fim, csg), iron acquisition (ent, fep, chu), motility, and secretion systems, with core virulence genes exhibiting > 90% sequence identity, whereas accessory virulence genes were more variable. Plasmid analysis demonstrated heterogeneous replicon types, predominantly IncF and Col plasmids, indicating their role in horizontal gene transfer. Jaccard similarity indices revealed moderate to high genetic overlap with global strains (~0.63 for virulence genes and ~0.55 for AMR profiles), suggesting shared evolutionary backgrounds. Phylogenomic and MLST identified all Bangladeshi isolates as ST457, clustering within a globally distributed clonal complex linked to ST10 and ST131 lineages. These findings suggest that the three Bangladeshi poultry-derived E. coli isolates are genetically related to globally circulating strains while harboring extensive resistance and virulence determinants, emphasizing poultry as an important reservoir of MDR pathogens and reinforcing the need for strengthened antimicrobial stewardship and genomic surveillance.

Animals↗

Single-read sequence tags of a limited number of genomic DNA fragments provide an inexpensive tool for comparative genome analysis.

Single-read sequences from both ends of 415 3-kb average size genomic DNA fragments of Candida albicans were compared with the complete sequence data of Saccharomyces cerevisiae. Comparison at the protein level, translated DNA against protein sequences, revealed 138 sequence tags with clear similarity to S. cerevisiae proteins or open reading frames. One case of synteny was found for the open reading frames of RAD16 and LYS2, which are adjacent to each other in S. cerevisiae and C. albicans.

Adenosine Triphosphatases↗