Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Comparative genomic analysis reveals distinct population structure in Legionella anisa.

Legionella anisa has been frequently isolated from engineered water systems; however, its population structure remains understudied compared to Legionella pneumophila. Here, we generated complete genome sequences for four L. anisa isolates recovered from a healthcare facility in Rimouski, Canada. Further the population structure of this species was investigated by performing comparative genomic analyses of the genomes generated in this study together with publicly available L. anisa genomes. Genome-wide phylogenetic analysis revealed the presence of three distinct clades separated by substantial genetic divergence (∼500 SNP), with the Rimouski isolates forming a tightly clustered group, suggesting a clonal lineage. Comparative pangenome analysis indicated moderate core genome conservation accompanied by a highly variable accessory genome (∼50%). The isolates characterized in this study harbored multiple plasmids encoding genes associated with conjugation, heavy metal resistance, and other stress-related functions, suggesting potential roles in environmental persistence. Previous studies have shown that L. anisa can proliferate within protozoan host cells, although outcomes vary depending on the host species. Our isolates showed efficient proliferation within Acanthamoeba castellanii, but not within Vermamoeba vermiformis, under the conditions tested. Together, these findings underscore the genomic diversity of this understudied Legionella species and provide a framework for future investigations regarding environmental persistence and potential pathogenicity.

Legionella anisa, Whole genome sequencing↗

GPAC: benchmarking the sensitivity of genome informatics analysis to genome annotation completeness.

In view of the recent explosion in genome sequence data, and the 200 or more complete genome sequences currently available, the importance of genome-scale bioinformatics analysis is increasing rapidly. However, computational genome informatics analyses often lack a statistical assessment of their sensitivity to the completeness of the functional annotation. Therefore, a pre-analysis method to automatically validate the sensitivity of computational genome analyses with regard to genome annotation completeness is useful for this purpose. In this report we developed the Gene Prediction Accuracy Classification (GPAC) test, which provides statistical evidence of sensitivity by repeating the same analysis for five different gene groups (classified according to annotation accuracy level), and for randomly sampled gene groups, with the same number of genes as each of the five classified groups. Variability in these results is then assessed, and if the results vary significantly with different data subsets, the analysis is considered "sensitive" to annotation completeness, and careful selection of data is advised prior to the actual in silico analysis. The GPAC test has been applied to the analyses of Sakai et al., 2001, and Ohno et al., 2001, and it revealed that the analysis of Ohno et al. was more sensitive to annotation completeness. It showed that GPAC could be employed to ascertain the sensitivity of an analysis. The GPAC bendhmarking software is freely available in the latest G-language Genome Analysis Environment package, at http://www.g-language.org/.

Benchmarking↗

Landscape genomics analysis reveals the genetic basis underlying cashmere goats and dairy goats adaptation to frigid environments.

Understanding the genetic mechanism of cold adaptation in cashmere goats and dairy goats is very important to improve their production performance. The purpose of this study was to comprehensively analyze the genetic basis of goat adaptation to cold environments, clarify the impact of environmental factors on genome diversity, and lay the foundation for breeding goat breeds to adapt to climate change. A total of 240 dairy goats were subjected to genome resequencing, and the whole genome sequencing data of 57 individuals from 6 published breeds were incorporated. By integrating multiple approaches such as phylogenetic analysis, population structure analysis, gene flow and population history exploration, selection signal analysis, and genome-environment association analysis, an in-depth investigation was carried out. Phylogenetic analysis unraveled the genetic relationships and differentiation patterns among dairy goats and other goat breeds. Through signal analysis (θπ, FST, XP-CLR), we identified numerous candidate genes associated with cold adaptation in dairy goats (STRIP1, ALX3, HTR4, NTRK2, MRPL11, PELI3, DPP3, BBS1) and cashmere goats (MED12L, MARC2, MARC1, DSG3, C6H4orf22, CHD7, MYPN, KIAA0825, MITF). Genome-environment association (GEA) analysis confirmed the link between these genes and environmental factors. Moreover, a detailed analysis of the critical genes C6H4orf22 and STRIP1 demonstrated their significant roles in the geographical variations of cold adaptation and allele frequency differences among different breeds. This study contributes to understanding the genetic basis of cold adaptation, providing crucial theoretical support for precision breeding programs aimed at improving production performance in cold regions by leveraging adaptive alleles, thereby ensuring sustainable animal husbandry.

Environmental adaptation↗

Complete genome analysis of the mandarin fish infectious spleen and kidney necrosis iridovirus.

The nucleotide sequence of the infectious spleen and kidney necrosis virus (ISKNV) genome was determined and found to comprise 111,362 bp with a G+C content of 54.78%. It contained 124 potential open reading frames (ORFs) with coding capacities ranging from 40 to 1208 amino acids. The analysis of the amino acid sequences deduced from the individual ORFs revealed that 35 of the 124 potential gene products of ISKNV show significant homology to functionally characterized proteins of other species. Some of the putative gene products of ISKNV showed significant homologies to proteins in the GenBank/EMBL/DDBJ databases including enzymes and structural proteins involved in virus replication, transcription, protein modification, and virus-host interaction. In addition, one major repeated sequence showing significant homology to the Red Sea bream iridovirus (RSIV) genome was identified. Based on the information obtained from biological properties (including histopathology, tissue tropisms, natural host range, and geographic distribution), physiochemical and physical properties, and genome analysis, we suggest that ISKNV, RSIV, sea bass iridovirus, grouper iridovirus, and African lampeye iridovirus may belong to a new genus of the Iridoviridae family and are tentatively referred to as cell hypertrophy iridoviruses.

Amino Acid Sequence↗

Comparative genome analysis of cortactin and HS1: the significance of the F-actin binding repeat domain.

BACKGROUND: In human carcinomas, overexpression of cortactin correlates with poor prognosis. Cortactin is an F-actin-binding protein involved in cytoskeletal rearrangements and cell migration by promoting actin-related protein (Arp)2/3 mediated actin polymerization. It shares a high amino acid sequence and structural similarity to hematopoietic lineage cell-specific protein 1 (HS1) although their functions differ considerable. In this manuscript we describe the genomic organization of these two genes in a variety of species by a combination of cloning and database searches. Based on our analysis, we predict the genesis of the actin-binding repeat domain during evolution. RESULTS: Cortactin homologues exist in sponges, worms, shrimps, insects, urochordates, fishes, amphibians, birds and mammalians, whereas HS1 exists in vertebrates only, suggesting that both genes have been derived from an ancestor cortactin gene by duplication. In agreement with this, comparative genome analysis revealed very similar exon-intron structures and sequence homologies, especially over the regions that encode the characteristic highly conserved F-actin-binding repeat domain. Cortactin splice variants affecting this F-actin-binding domain were identified not only in mammalians, but also in amphibians, fishes and birds. In mammalians, cortactin is ubiquitously expressed except in hematopoietic cells, whereas HS1 is mainly expressed in hematopoietic cells. In accordance with their distinct tissue specificity, the putative promoter region of cortactin is different from HS1. CONCLUSIONS: Comparative analysis of the genomic organization and amino acid sequences of cortactin and HS1 provides inside into their origin and evolution. Our analysis shows that both genes originated from a gene duplication event and subsequently HS1 lost two repeats, whereas cortactin gained one repeat. Our analysis genetically underscores the significance of the F-actin binding domain in cytoskeletal remodeling, which is of importance for the major role of HS1 in apoptosis and for cortactin in cell migration.

Actin-Related Protein 2↗

[Ras gene analysis in mammary tumors of dogs by means of PCR-SSCP and direct genomic analysis].

The oncogenic capacities of RAS family genes (Ha-ras, Ki-ras, and N-ras) are usually activated by point mutations in the conserved regions (codons 12, 13, and 61), resulting in single amino acid substitution in the specific proteins (p21). In order to verify the involvement of RAS genes in dog mammary tumors we analyzed the genomic DNA from 20 mammary tumors of dog by means of the Polymerase Chain Reaction-Single Strand Conformation Polymorphism (PCR-SSCP) method and the direct genomic sequencing. The absence of point mutations in the "hot spots" of RAS genes suggests a lack or a low frequency of such a pattern of RAS genes activation in dog mammary tumors. The results are also in agreement to what reported in human mammary tumors. However, the presence of genetic alterations in other functional areas of the RAS genes or other mechanisms of activations cannot be ruled out.

Animals↗

Hundredfold productivity of genome analysis by introduction of microtemperature-gradient gel electrophoresis.

Genome profiling, which employs temperature-gradient gel electrophoresis (TGGE) for DNA analysis, has recently been developed in identifying species by genotype. However, the performance of this technology like the general applications of TGGE was, though highly informative, limited in its ability due to methodological reasons. This study demonstrates that minimization of the gel for TGGE, to around one-tenth of its conventional size (approximately 2 cm), can be successfully introduced, resulting in a hundredfold higher performance (total evaluation of time, cost, and degree of parallel operations) than that of the conventional. Reproducibility was evaluated from the measures of the pattern similarity scores (PaSS) between band patterns (genome profiles) obtained with the conventional TGGE, and that with micro-TGGE (microTGGE) developed here, after extracting a set of featuring points from genome profiles. Size minimization, which leads to the reduction of the amount of samples required (cost-saving), is another great advantage, enhancing the employment of multicolor fluorescence technology. Since the further development of microbe-related fields such as epidemiology and microbial ecology inevitably require knowledge based on the identification of a great number of species and strains, microbe-related fields will receive the most optimal benefits from the technological improvements attained here.

Bacillus↗

Genome analysis of the obligately lytic bacteriophage 4268 of Lactococcus lactis provides insight into its adaptable nature.

Analysis of the complete nucleotide sequence of the lactococcal phage 4268, which is lytic for the cheese starter Lactococcus lactis DPC4268, is presented. Phage 4268 has a linear genome of 36,596 bp, which is modularly organised and encompasses 49 open reading frames. Putative functions were assigned to approximately 45% of the predicted products of these open reading frames based on sequence similarity with known proteins, N-terminal sequence analysis and identification of conserved domains. Significantly, a segment of the genome has homology to the recently sequenced lysogenic module in lactococcal phage phi31 that contains a lytic switch but no phage integrase or attachment site. This suggests that it is derived from a prophage. A phage 4268-encoded and a host-encoded methylase were found to be highly similar, having only two nucleotide mismatches, suggesting that the phage acquired the methylase gene to protect it from a host endonuclease. Comparative genomic analysis revealed significant homology between phage 4268 and the lactococcal phage BK5-T. The comparative analysis also supported the classification of phage 4268 and other BK5-T-related phage as separate from the proposed P335 species of lactococcal phage.

Bacteriophages↗

New approaches in genome analysis by pulsed-field gel electrophoresis: application to the analysis of Pseudomonas species.

A general method for the evaluation of macrorestriction fragment patterns is presented and its applicability to the taxonomy of bacteria is demonstrated for 32 Pseudomonas species. Strains were differentiated at the species and subspecies level by genome size and macrorestriction fragment fingerprints of the chromosome that had been separated on pulsed-field gels. The relatedness of bacteria was ascertained from the similarity of AsnI, DraI, SpeI, SspI or XbaI fragment patterns. In general, the dendrograms calculated from the genome fingerprints corresponded with the phylogenetic classification obtained from phenotypic marker or nucleic acid hybridization analysis, but several exceptions were noted. The techniques and algorithms presented herein are generally applicable to the genome analysis of bacteria, lower eukaryotes, and DNA fragments cloned in yeast artificial chromosomes.

Bacteria↗

Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.

Determining protein functions from genomic sequences is a central goal of bioinformatics. We present a method based on the assumption that proteins that function together in a pathway or structural complex are likely to evolve in a correlated fashion. During evolution, all such functionally linked proteins tend to be either preserved or eliminated in a new species. We describe this property of correlated evolution by characterizing each protein by its phylogenetic profile, a string that encodes the presence or absence of a protein in every known genome. We show that proteins having matching or similar profiles strongly tend to be functionally linked. This method of phylogenetic profiling allows us to predict the function of uncharacterized proteins.

Bacterial Proteins↗

Identification of novel virulence-associated genes via genome analysis of hypothetical genes.

The sequencing of bacterial genomes has opened new perspectives for identification of targets for treatment of infectious diseases. We have identified a set of novel virulence-associated genes (vag genes) by comparing the genome sequences of six human pathogens that are known to cause persistent or chronic infections in humans: Yersinia pestis, Neisseria gonorrhoeae, Helicobacter pylori, Borrelia burgdorferi, Streptococcus pneumoniae, and Treponema pallidum. This comparison was limited to genes annotated as hypothetical in the T. pallidum genome project. Seventeen genes with unknown functions were found to be conserved among these pathogens. Insertional inactivation of 14 of these genes generated nine mutants that were attenuated for virulence in a mouse infection model. Out of these nine genes, five were found to be specifically associated with virulence in mice as demonstrated by infection with Yersinia pseudotuberculosis in-frame deletion mutants. In addition, these five vag genes were essential only in vivo, since all the mutants were able to grow in vitro. These genes are broadly conserved among bacteria. Therefore, we propose that the corresponding vag gene products may constitute novel targets for antimicrobial therapy and that some vag mutants could serve as carrier strains for live vaccines.

Animals↗

Comparative genomic analysis of hyperthermophilic archaeal Fuselloviridae viruses.

The complete genome sequences of two Sulfolobus spindle-shaped viruses (SSVs) from acidic hot springs in Kamchatka (Russia) and Yellowstone National Park (United States) have been determined. These nonlytic temperate viruses were isolated from hyperthermophilic Sulfolobus hosts, and both viruses share the spindle-shaped morphology characteristic of the Fuselloviridae family. These two genomes, in combination with the previously determined SSV1 genome from Japan and the SSV2 genome from Iceland, have allowed us to carry out a phylogenetic comparison of these geographically distributed hyperthermal viruses. Each virus contains a circular double-stranded DNA genome of approximately 15 kbp with approximately 34 open reading frames (ORFs). These Fusellovirus ORFs show little or no similarity to genes in the public databases. In contrast, 18 ORFs are common to all four isolates and may represent the minimal gene set defining this viral group. In general, ORFs on one half of the genome are colinear and highly conserved, while ORFs on the other half are not. One shared ORF among all four genomes is an integrase of the tyrosine recombinase family. All four viral genomes integrate into their host tRNA genes. The specific tRNA gene used for integration varies, and one genome integrates into multiple loci. Several unique ORFs are found in the genome of each isolate.

Archaeal Viruses↗

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics↗

[Genome analysis of pathogenic bacteria--a review].

The whole genome information on pathogenic bacteria opened a new era of medical microbiology. Bacteria have a variety of strategies to survive in human body against host defense. Bacteria colonizing on the mucous cell surface and proliferate extracellularly have developed a number of paralogs of surface proteins for antigenic variation, while bacteria which proliferate intracellularly have acquired genes for invasion and anti-active oxygen. The obligate intracellular parasites, Rickettsiae and Chlamydiae, have developed a distinct energy metabolism depending on the intracellular localization; Rickettsia prowazekii has TCA cycle and oxidative phosphorylation to proliferate in the cytoplasm, while Chlamydiae have a glycolytic pathway and substrate-level phosphorylation to multiply within an organella called inclusion bodies.

Antigenic Variation↗

Antigen expression from cloned genes of Mycoplasma hyorhinis: an approach to mycoplasma genomic analysis.

Cloned DNA segments of the Mycoplasma hyorhinis genome have been identified by specific immunological detection of gene products expressed in Escherichia coli. Individual, distinct segments collectively representing greater than 10% of the genome each synthesize multiple proteins, some of which are immunogenic. Molecular genetic and immunologic tools have identified mycoplasma genomic fragments useful in 1) generating and analyzing specific, isolated mycoplasma gene products, 2) studying structure and expression of specific mycoplasma genes, and 3) generating specific gene probes. Immunological detection provides the potential for preselection of specific genes of interest.

Antigens, Bacterial↗

Whole genome analysis of a multidrug-resistant blaNDM-5-carrying Escherichia coli Sequence Type (ST) 167 strain isolated from seafood in Mumbai, India.

BACKGROUND: E. coli ST167 is an emerging extraintestinal pathogenic Escherichia coli (ExPEC) clone. This study reports the whole genome sequence analysis of a multidrug-resistant, blaNDM-5 harboring E. coli ST167 (EC121) isolated from seafood. The antibiotic susceptibility pattern was determined using the standard disc diffusion method. Genomic DNA was extracted, purified, and sequenced using the Illumina platform. The whole genome sequence was analyzed to determine the genome characteristics, including sequence type, serotype, phylogroup, antibiotic resistance genes, virulence attributes, and phylogenetic analysis. RESULTS: Phenotypically, this isolate was resistant to 26 of the 33 antibiotics tested, which correlated well with in-silico prediction. Multilocus sequence typing (MLST) analysis revealed that this strain belonged to sequence type 167, serotype O101:H9, and phylogroup A and harbored different virulence genes, suggesting it was a potential human pathogen. Many acquired antibiotic resistance genes were detected, including blaNDM-5, blaCMY-42, blaOXA-1, blaTEM-116, catA1, sul2, and tet(B). Point mutations in gyrA and parC responsible for quinolone resistance were also detected. CONCLUSION: The combinations of virulence and antibiotic resistance genes in this strain highlight the significant risk associated with emerging E. coli clonal types contaminating the seafood supply chain. Fecal contamination of seafood can contribute to the community dissemination of multidrug-resistant E. coli, necessitating effective monitoring measures.

Seafood↗

Genomic analysis of avian and feline ureaplasmas by restriction endonucleases.

Genomic relationship among the strains of avian and feline ureaplasmas was determined by restriction analysis. Chromosomal DNAs extracted from the two avian and four feline ureaplasma strains resolved in 1.0% agarose gel electrophoresis following digestion with the restriction enzymes, PstI, BamHI, HindIII, EcoRI, SalI and HpaII. Relationship among the strains was determined from base substitution frequency obtained by comparing the restriction patterns. The guanine plus cytosine (G + C) contents of the representative strains of avian and feline ureaplasmas were estimated by high-performance liquid chromatography to be 27.6 and 27.1 mole %, respectively. Restriction patterns of the avian and feline ureaplasmas were distinct from those of the human and bovine strains, but they were relatively similar within a serotype.

Animals↗

Lineage dynamics of invasive Escherichia coli isolates in the Netherlands from 1975 to 2021: a retrospective longitudinal genomic analysis.

BACKGROUND: Escherichia coli is a common cause of invasive infections such as bloodstream and cerebrospinal fluid infections in neonates. Strains positive for the K1 capsule are considered the most common cause of such neonatal invasive infections. This assumption of K1 dominance, and indeed the population genomics of E coli causing invasive infections in general is largely unstudied. We aimed to provide a comprehensive characterisation of this pathogen population using a longitudinal isolate collection. METHODS: In this analysis we report the findings of the SENTINEL study, a longitudinal genomic analysis of 1790 invasive E coli isolates collected mainly from newborns in the Netherlands between 1975 and 2021 by the Netherlands Reference Laboratory for Bacterial Meningitis, Amsterdam University Medical Centre, Amsterdam, Netherlands. The dataset included all bacterial strains cultured from cerebrospinal fluid or blood in cases of (clinical) bacterial meningitis (1976 to 1980). In 1981 the criteria were expanded to include neonates (aged ≤4 weeks) with E coli sepsis, and from July, 2016 all infants younger than 1 year with E coli sepsis were included. All isolates were sequenced using either the HiSeq 2500 or HiSeq 4000 platforms (Illumina, San Diego, CA, USA). We confirmed species and identified sequence types (STs), detected antimicrobial resistance genes, virulence genes, and the presence of K1 capsule, and characterised the dynamics of these factors over time. FINDINGS: Our data show a highly dynamic bacterial population that is entirely unaffected by antimicrobial resistance determinants. Key pathogen population fluctuations include the complete disappearance of the dominant lineage ST567 and the swapping of dominant ST95 clones from a single serotype O18:H7 clone to two distinct serotype O1:H7 clones, with changes in virulence factors including major fimbrial adhesins. These findings, combined with only 58·8% (1053 of 1790) prevalence in K1-expressing isolates in the entire study population, point to host-pathogen interaction and immune selection pressures as key drivers of bacterial population dynamics in this largely antimicrobial-naive population. INTERPRETATION: Our data show the vital need for ongoing genomic surveillance of microbial pathogen populations to guide appropriate intervention strategies. Additionally, genomic insights of a pathogen population from one specific disease syndrome or patient population cannot always be generalised across other cohorts. FUNDING: Wellcome Antimicrobial and Antimicrobial Resistance Doctoral Training Programme and the National Institute for Health and Care Research Birmingham Biomedical Research Centre.

Netherlands↗