Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comparative genomic analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Large-scale genetic variation of the symbiosis-required megaplasmid pSymA revealed by comparative genomic analysis of Sinorhizobium meliloti natural strains.

BACKGROUND: Sinorhizobium meliloti is a soil bacterium that forms nitrogen-fixing nodules on the roots of leguminous plants such as alfalfa (Medicago sativa). This species occupies different ecological niches, being present as a free-living soil bacterium and as a symbiont of plant root nodules. The genome of the type strain Rm 1021 contains one chromosome and two megaplasmids for a total genome size of 6 Mb. We applied comparative genomic hybridisation (CGH) on an oligonucleotide microarrays to estimate genetic variation at the genomic level in four natural strains, two isolated from Italian agricultural soil and two from desert soil in the Aral Sea region. RESULTS: From 4.6 to 5.7 percent of the genes showed a pattern of hybridisation concordant with deletion, nucleotide divergence or ORF duplication when compared to the type strain Rm 1021. A large number of these polymorphisms were confirmed by sequencing and Southern blot. A statistically significant fraction of these variable genes was found on the pSymA megaplasmid and grouped in clusters. These variable genes were found to be mainly transposases or genes with unknown function. CONCLUSION: The obtained results allow to conclude that the symbiosis-required megaplasmid pSymA can be considered the major hot-spot for intra-specific differentiation in S. meliloti.

Blotting, Southern↗

Comparative genome analysis of the neurexin gene family in Danio rerio: insights into their functions and evolution.

Neurexins constitute a family of proteins originally identified as synaptic transmembrane receptors for a spider venom toxin. In mammals, the 3 known Neurexin genes present 2 alternative promoters that drive the synthesis of a long (alpha) and a short (beta) form and contain different sites of alternative splicing (AS) that can give rise to thousands of different transcripts. To date, very little is known about the significance of this variability, except for the modulation of binding to some of the Neurexin ligands. Although orthologs of Neurexins have been isolated in invertebrates, these genes have been studied mostly in mammals. With the aim of investigating their functions in lower vertebrates, we chose Danio rerio as a model because of its increasing importance in comparative biology. We have isolated 6 zebrafish homologous genes, which are highly conserved at the structural level and display a similar regulation of AS, despite about 450 Myr separating the human and zebrafish species. Our data indicate a strong selective pressure at the exonic level and on the intronic borders, in particular on the regulative intronic sequences that flank the exons subject to AS. Such a selective pressure could help conserve the regulation and consequently the function of these genes along the vertebrates evolutive tree. AS analysis during development shows that all genes are expressed and finely regulated since the earliest stages of development, but mark an increase after the 24-h stage that corresponds to the beginning of synaptogenesis. Moreover, we found that specific isoforms of a zebrafish Neurexin gene (nrxn1a) are expressed in the adult testis and in the earliest stages of development, before the beginning of zygotic transcription, indicating a potential delivery of paternal RNA to the embryo. Our analysis suggests the existence of possible new functions for Neurexins, serving as the basis for novel approaches to the functional studies of this complex neuronal protein family and more in general to the understanding of the AS mechanism in low vertebrates.

Alternative Splicing↗

The fmo genes of Caenorhabditis elegans and C. briggsae: characterisation, gene expression and comparative genomic analysis.

The flavin-containing monooxygenase (FMO) gene family is conserved and ancient with representatives present in almost all phyla so far examined. The genes encode FAD-, NADP- and O(2)-dependent enzymes that catalyse oxygenation of soft-nucleophilic heteroatom centres in a range of substrates. Although usually classified as xenobiotic-metabolising enzymes, examples of FMOs exist that have evolved to metabolise specific endogenous substrates as part of a discrete physiological process. The genome of Caenorhabditis elegans contains five predicted genes encoding putative homologs of mammalian FMOs, K08C7.2, K08C7.5, Y39A1A.19, F53F4.5 and H24K24.5, which we have named fmo and numbered fmo-1 to fmo-5, respectively. As a first step towards determining their functional role(s), we have experimentally characterised these C. elegans fmo genes including analysing reporter gene expression patterns and RNAi phenotypes. Two major gene expression patterns were observed, either intestinal or hypodermal, but no gross RNAi phenotypes were found possibly due to functional redundancy. The internal structures of fmo-2, fmo-3 and fmo-4 have been compared with orthologs identified in the related nematode C. briggsae. For each orthologous pair, a global comparison of the paired upstream intergenic regions was performed and a number of conserved noncoding sequences, which may represent potential cis-regulatory elements, identified. Phylogenetic analysis reveals that several of the fmo homologs are the result of gene duplication along the lineage leading to the nematodes.

Amino Acid Sequence↗

A comparative genomic analysis of two distant diptera, the fruit fly, Drosophila melanogaster, and the malaria mosquito, Anopheles gambiae.

Genome evolution entails changes in the DNA sequence of genes and intergenic regions, changes in gene numbers, and also changes in gene order along the chromosomes. Genes are reshuffled by chromosomal rearrangements such as deletions/insertions, inversions, translocations, and transpositions. Here we report a comparative study of genome organization in the main African malaria vector, Anopheles gambiae, relative to the recently determined sequence of the Drosophila melanogaster genome. The ancestral lines of these two dipteran insects are thought to have separated approximately 250 Myr, a long period that makes this genome comparison especially interesting. Sequence comparisons have identified 113 pairs of putative orthologs of the two species. Chromosomal mapping of orthologous genes reveals that each polytene chromosome arm has a homolog in the other species. Between 41% and 73% of the known orthologous genes remain linked in the respective homologous chromosomal arms, with the remainder translocated to various nonhomologous arms. Within homologous arms, gene order is extensively reshuffled, but a limited degree of conserved local synteny (microsynteny) can be recognized.

Animals↗

Comparative genomic analysis of a novel heat-tolerant and euryhaline strain of unicellular marine cyanobacterium Cyanobacterium sp. DS4 from a high-temperature lagoon.

BACKGROUND: Cyanobacteria have diversified through their long evolutionary history and occupy a wide range of environments on Earth. To advance our understanding of their adaptation mechanisms in extreme environments, we performed stress tolerance characterizations, whole genome sequencing, and comparative genomic analyses of a novel heat-tolerant and euryhaline strain of the unicellular cyanobacterium Cyanobacterium sp. Dongsha4 (DS4). This strain was isolated from a lagoon on Dongsha Island in the South China Sea, a habitat with fluctuations in temperature, salinity, light intensity, and nutrient supply. RESULTS: DS4 cells can tolerate long-term high-temperature up to 50 ℃ and salinity from 0 to 6.6%, which is similar to the results previously obtained for Cyanobacterium aponinum. In contrast, most mesophilic cyanobacteria cannot survive under these extreme conditions. Based on the 16S rRNA gene phylogeny, DS4 is most closely related to Cyanobacterium sp. NBRC102756 isolated from Iwojima Island, Japan, and Cyanobacterium sp. MCCB114 isolated from Vypeen Island, India. For comparison with strains that have genomic information available, DS4 is most similar to Cyanobacterium aponinum strain PCC10605 (PCC10605), sharing 81.7% of the genomic segments and 92.9% average nucleotide identity (ANI). Gene content comparisons identified multiple distinct features of DS4. Unlike related strains, DS4 possesses the genes necessary for nitrogen fixation. Other notable genes include those involved in photosynthesis, central metabolisms, cyanobacterial starch metabolisms, stress tolerances, and biosynthesis of novel secondary metabolites. CONCLUSIONS: These findings promote our understanding of the physiology, ecology, evolution, and stress tolerance mechanisms of cyanobacteria. The information is valuable for future functional studies and biotechnology applications of heat-tolerant and euryhaline marine cyanobacteria.

Cyanobacteria↗

Comparative genomic analysis, diversity and evolution of two KIR haplotypes A and B.

Members of the killer immunoglobulin (Ig)-like receptor (KIR) gene family are tightly clustered on human chromosome 19q13.4. Despite considerable variation in KIR gene content and allelic polymorphism, most KIR haplotypes belong to one of two broad groups termed A and B. The availability of contiguous genomic sequences for these haplotypes has allowed us to compare their genomic organization, nucleotide (nt) diversity and reconstruct their evolutionary history. The haplotypes have a framework of three conserved blocks containing (i) KIR3DL3, (ii) KIR3DP1, 2DL4, and (iii) KIR3DL2 that are interrupted by two variable segments that differ in the number and type of KIR genes. Low (0.05%) nucleotide diversity was detected across the centromeric and telomeric boundaries of the KIR gene cluster while higher SNP density (0.2%) occurred within the central region containing the KIR2DL4 gene. Phylogenetic and genomic analyses have permitted the reconstruction of a hypothetical ancestral haplotype that has revealed common groupings and differences between the KIR genes of the two haplotypes. The present phylogenetic and genomic comparison of the two sequenced KIR haplotypes provides a framework for a more thorough examination of KIR haplotype variations, diversity and evolution in human populations and between humans and non-human primates.

Animals↗

Comparative genome analysis reveals extensive conservation of genome organisation for Arabidopsis thaliana and Capsella rubella.

Genome colinearity has been studied for two closely related diploid species of the Brassicaceae family, Arabidopsis thaliana and Capsella rubella. Markers mapping to chromosome 4 of A. thaliana were found on two linkage groups in Capsella and colinear segments spanning more than 10 cM were revealed. Detailed analysis of a 60 kbp region in A. thaliana and its counterpart in C. rubella showed virtually complete conservation of gene repertoire, order and orientation. The comparison of orthologous genes revealed very similar exon-intron structures and sequence identities of 90% or more were found for exon sequences. This extensive genome colinearity at the genetic and molecular level allows the efficient transfer of data from the well-studied A. thaliana genome to other species in the Brassicaceae family, substantially facilitating genome analysis studies for species of this family.

Amino Acid Sequence↗

Comparative genomic analysis of key oncogenic pathways in hepatocellular carcinoma among diverse populations.

BACKGROUND/OBJECTIVES: Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality, with significant racial and ethnic disparities in incidence, tumor biology, and clinical outcomes. Hispanic/Latino (H/L) patients tend to be diagnosed at younger ages and more advanced stages than Non-Hispanic White (NHW) patients, yet the molecular mechanisms underlying these disparities remain poorly understood. Key oncogenic pathways, including RTK/RAS, TGF-Beta, WNT, PI3K, and TP53, play pivotal roles in tumor progression, treatment resistance, and response to targeted therapies. However, ethnicity-specific alterations within these pathways remain largely unexplored. This study aims to compare pathway-specific mutations in HCC between H/L and NHW patients, assess tumor mutation burden, and identify ethnicity-associated oncogenic drivers using publicly available datasets. Findings from this analysis may inform precision medicine strategies for improving early detection and targeted therapies in underrepresented populations. METHODS: We conducted a bioinformatics analysis using publicly available HCC datasets to assess mutation frequencies in RTK/RAS, TGF-Beta, WNT, PI3K, and TP53 pathway genes. The study included 547 patients, consisting of 69 H/L patients and 478 NHW patients. Patients were stratified by ethnicity (H/L vs. NHW) to evaluate differences in mutation prevalence. Chi-squared tests were used to compare mutation frequencies, while Kaplan-Meier survival analysis assessed overall survival differences associated with pathway-specific alterations in both populations. RESULTS: Significant differences were observed in the RTK/RAS pathway related genes, particularly in FGFR4 mutations, which were more prevalent in H/L patients compared to NHW patients (4.3% vs. 0.6%, p = 0.02). Additionally, IGF1R mutations exhibited borderline significance (7.2% vs. 2.9%, p = 0.07). In the PI3K pathway, INPP4B alterations were more frequent in H/L patients than in NHW patients (4.3% vs. 1%, p = 0.06), while in the TGF-Beta pathway, TGFBR2 mutations were more common in H/L patients (2.9% vs. 0.4%, p = 0.07), suggesting potential ethnicity-specific variations. Survival analysis revealed no significant differences in overall survival between H/L and NHW patients, indicating that molecular alterations alone may not fully explain survival disparities and suggesting a role for additional factors such as immune response, environmental exposures, or access to targeted therapies. CONCLUSIONS: This study provides one of the first ethnicity-focused analyses of key oncogenic pathway alterations in HCC, revealing distinct molecular differences between H/L and NHW patients. The findings suggest that RTK/RAS (FGFR4, IGF1R), PI3K (INPP4B), and TGF-Beta (TGFBR2) pathway alterations may play a distinct role in HCC among H/L patients, while their prognostic significance in NHW patients remains unclear. These insights emphasize the importance of incorporating ethnicity-specific molecular profiling into precision medicine approaches to improve early detection, targeted therapies, and clinical outcomes in HCC, particularly for underrepresented populations.

PI3K pathway↗

Identification of conserved regulatory elements by comparative genome analysis.

BACKGROUND: For genes that have been successfully delineated within the human genome sequence, most regulatory sequences remain to be elucidated. The annotation and interpretation process requires additional data resources and significant improvements in computational methods for the detection of regulatory regions. One approach of growing popularity is based on the preferential conservation of functional sequences over the course of evolution by selective pressure, termed 'phylogenetic footprinting'. Mutations are more likely to be disruptive if they appear in functional sites, resulting in a measurable difference in evolution rates between functional and non-functional genomic segments. RESULTS: We have devised a flexible suite of methods for the identification and visualization of conserved transcription-factor-binding sites. The system reports those putative transcription-factor-binding sites that are both situated in conserved regions and located as pairs of sites in equivalent positions in alignments between two orthologous sequences. An underlying collection of metazoan transcription-factor-binding profiles was assembled to facilitate the study. This approach results in a significant improvement in the detection of transcription-factor-binding sites because of an increased signal-to-noise ratio, as demonstrated with two sets of promoter sequences. The method is implemented as a graphical web application, ConSite, which is at the disposal of the scientific community at http://www.phylofoot.org/. CONCLUSIONS: Phylogenetic footprinting dramatically improves the predictive selectivity of bioinformatic approaches to the analysis of promoter sequences. ConSite delivers unparalleled performance using a novel database of high-quality binding models for metazoan transcription factors. With a dynamic interface, this bioinformatics tool provides broad access to promoter analysis with phylogenetic footprinting.

Algorithms↗

Comparative genomic analysis as a tool for biological discovery.

The recent completion of the human genome sequence has enabled the identification of a large fraction of our gene catalogue and their physical chromosomal position. However, current efforts lag at defining the cis-regulatory sequences that control the spatial and temporal patterns of each gene's expression. This task remains difficult due to our lack of knowledge of the vocabulary controlling gene regulation and the vast genomic search space, with greater than 95% of our genome being noncoding. Recent comparative genomic-based strategies are beginning to aid in the identification of functional sequences based on their high levels of evolutionary conservation. This has proven successful for comparisons between closely related species such as human-primate or human-mouse, but also holds true for distant evolutionary comparisons, such as human-fish or human-bird. In this review we provide support for the utility of cross-species sequence comparisons by illustrating several applications of this strategy, including the identification of new genes and functional non-coding sequences. We also discuss emerging concepts as this field matures, such as how to properly select which species for comparison, which may differ significantly between independent studies.

Animals↗

Comparative genome analysis reveals a conserved family of actin-like proteins in apicomplexan parasites.

BACKGROUND: The phylum Apicomplexa is an early-branching eukaryotic lineage that contains a number of important human and animal pathogens. Their complex life cycles and unique cytoskeletal features distinguish them from other model eukaryotes. Apicomplexans rely on actin-based motility for cell invasion, yet the regulation of this system remains largely unknown. Consequently, we focused our efforts on identifying actin-related proteins in the recently completed genomes of Toxoplasma gondii, Plasmodium spp., Cryptosporidium spp., and Theileria spp. RESULTS: Comparative genomic and phylogenetic studies of apicomplexan genomes reveals that most contain only a single conventional actin and yet they each have 8-10 additional actin-related proteins. Among these are a highly conserved Arp1 protein (likely part of a conserved dynactin complex), and Arp4 and Arp6 homologues (subunits of the chromatin-remodeling machinery). In contrast, apicomplexans lack canonical Arp2 or Arp3 proteins, suggesting they lost the Arp2/3 actin polymerization complex on their evolutionary path towards intracellular parasitism. Seven of these actin-like proteins (ALPs) are novel to apicomplexans. They show no phylogenetic associations to the known Arp groups and likely serve functions specific to this important group of intracellular parasites. CONCLUSION: The large diversity of actin-like proteins in apicomplexans suggests that the actin protein family has diverged to fulfill various roles in the unique biology of intracellular parasites. Conserved Arps likely participate in vesicular transport and gene expression, while apicomplexan-specific ALPs may control unique biological traits such as actin-based gliding motility.

Actins↗

Bacterial artificial chromosome-based comparative genomic analysis identifies Mycobacterium microti as a natural ESAT-6 deletion mutant.

Mycobacterium microti is a member of the Mycobacterium tuberculosis complex that causes tuberculosis in voles. Most strains of M. microti are harmless for humans, and some have been successfully used as live tuberculosis vaccines. In an attempt to identify putative virulence factors of the tubercle bacilli, genes that are absent from the avirulent M. microti but present in human pathogen M. tuberculosis or Mycobacterium bovis were searched for. A minimal set of 50 bacterial artificial chromosome (BAC) clones that covers almost all of the genome of M. microti OV254 was constructed, and individual BACs were compared to the corresponding BACs from M. bovis AF2122/97 and M. tuberculosis H37Rv. Comparison of pulsed-field gel-separated DNA digests of BAC clones led to the identification of 10 regions of difference (RD) between M. microti OV254 and M. tuberculosis. A 14-kb chromosomal region (RD1(mic)) that partly overlaps the RD1 deletion in the BCG vaccine strain was missing from the genomes of all nine tested M. microti strains. This region covers 13 genes, Rv3864 to Rv3876, in M. tuberculosis, including those encoding the potent ESAT-6 and CFP-10 antigens. In contrast, RD5(mic), a region that contains three phospholipase C genes (plcA to -C), was missing from only the vole isolates and was present in M. microti strains isolated from humans. Apart from RD1(mic) and RD5(mic) other M. microti-specific deleted regions have been identified (MiD1 to MiD3). Deletion of MiD1 has removed parts of the direct repeat region in M. microti and thus contributes to the characteristic spoligotype of M. microti strains.

Animals↗

Comparative genomic analysis of genes encoding translation elongation factor 1B(alpha) in human and mouse shows EEF1B1 to be a recent retrotransposition event.

We have characterized genomic loci encoding translation elongation factor 1B(alpha) (eEF1B(alpha)) in mice and humans. Mice have a single structural locus (named Eef1b2) spanning six exons, which is ubiquitously expressed and maps close to Casp8 on mouse chromosome 1, and a processed pseudogene. Humans have a single intron-containing locus, EEF1B2, which maps to 2q33, and an intronless paralogue expressed only in brain and muscle (EEF1B3). Another locus described previously, EEF1B1, is actually a processed pseudogene on chromosome 15 corresponding to an alternative splice form of EEF1B2. Our study illustrates the value of comparative mapping in distinguishing between processed pseudogenes and intronless paralogues.

Alternative Splicing↗

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

Characterization and comparative genomic analysis of intronless Adams with testicular gene expression.

ADAM (a disintegrin and metalloprotease) family members with testis-specific or -predominant gene expression are divided phylogenically into two groups: ADAMs 2, 3, 5, 27, and 32 (the first group) and ADAMs 4, 6, 20, 21, 24, 25, 26, 29, 30, and 34 (the second group). We cloned and sequenced cDNAs for previously unidentified mouse Adams that belong to the second group. We found that all the Adam genes in the second phylogenic group are transcribed by both somatic and germ cells in mouse testis, representing a unique expression pattern different from that of the first-group Adams. Genomic analyses revealed that all the second-group Adam genes lack introns interrupting protein-coding sequences and many of them are present as multicopy genes, resulting in total of 14 functional mouse genes in this phylogenic group. Comparing the mouse and human ADAM genes, we found that a number of these mouse Adam genes do not have human orthologues and, even if they exist, some orthologues are pseudogenes in human. These results suggest the differential expansion of the second-group Adam genes in the mouse genome during evolution and a relationship between these Adams and male reproduction unique to mouse.

Amino Acid Sequence↗