Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Deciphering transcriptional regulatory elements that encode specific cell cycle phasing by comparative genomics analysis.

Transcriptional regulation is a major tier in the periodic engine that mobilizes cell cycle progression. The availability of complete genome sequences of multiple organisms holds promise for significantly improving the specificity of computational identification of functional elements. Here, we applied a comparative genomics analysis to decipher transcriptional regulatory elements that control cell cycle phasing. We analyzed genome-wide promoter sequences from 12 organisms, including worm, fly, fish, rodents and human, and identified conserved transcriptional modules that determine the expression of genes in specific cell cycle phases. We demonstrate that a canonical E2F signal encodes for expression highly specific to the G1/S phase, and that a cis-regulatory module comprising CHR-NF-Y elements dictates expression that is restricted to the G2 and G2/M phases. B-Myb binding site signatures occur in many of the CHR-NF-Y target genes, suggesting a specific role for this triplet in the regulation of the cell cycle transcriptional program. Remarkably, E2F signals are conserved in promoters of G1/S genes in all organisms from worm to human. The CHR-NF-Y module is conserved in promoters of G2/M regulated genes in all analyzed vertebrates. Our results reveal novel modules that determine specific cell cycle phasing, and identify their respective putative target genes with remarkably high specificity.

Animals↗

[Genome analysis and primer sets determination for listeria species detection and gene typing].

Computer analysis of listeria genome isolates of sequences from EMBL, GenBank, DDBJ data bases has been made. Variable and highly conservative (homology degree is 90-100% for all known isolates) genes loci iap and listeriolysin (cytolysin) gene locuses have been determined. Primer sets for detection and differentiation of Listeria species by polymerase chain reaction (PCR) were designed by computer and following thermodynamic analysis. Primer sets can provide detection of Listeria monocytogenes, L. ivanovii, L. seeligeri, L. grayi, L. innocua, L. welshimeri and detection of only pathogenic L. monocytogenes and Listeria species with listeriolysin (cytolysin) gene. Differentiation of 6 Listeria species can be done by means of two primer sets for iap and listeriolysin genes. Amplified fragments of listeria species DNA have different amplicon size. Primers melting temperatures selection allows one to carry out listeria species typing by multiplex PCR in a single tube.

Bacterial Proteins↗

Genomic analysis reveals a duplication of eight rather than seven short consensus repeats in primate CR1 and CR1L: evidence for an additional set shared between CR1 and CR2.

We report the discovery of previously unrecognised short consensus repeats (SCRs) within human and chimpanzee CR1 and CR1L. Analysis of available genomic, protein and expression databases suggests that these are actually genomic remnants of SCRs previously reported in other complement control proteins (CCPs). Comparison with the nucleotide motifs of the 11 defined subfamilies of SCRs justifies the designation g-like because of the close similarity to the g subfamily found in CR2 and MCP. To date, we have identified five such SCRs in human and chimpanzee CR1, one in human and chimpanzee CR1L, but none in either rat or mouse Crry in keeping with the number of internal duplications of the long homologous repeat (LHR) found in CR1 and CR1L. In fact, at the genomic level, the ancestral LHR must have contained eight SCRs rather than seven as previously thought. Since g-like SCRs are found immediately downstream of d SCRs, we suggest that there must have been a functional dg set which has been retained by CR2 and MCP but which is degenerate in CR1 or CR1L. Interestingly, dg is also present in the CR2 component of mouse CR1. The degeneration of the g SCR must have occurred prior to the formation of primate CR1L and prior to the duplication events which resulted in primate CR1. In this context, the apparent conservation of g-like SCRs may be surprising and may suggest the existence of mechanisms unrelated to protein coding. These results provide examples of the many processes which have contributed to the evolution of the extensive repertoire of CCPs.

Animals↗

4th International Meeting on Single Nucleotide Polymorphism and Complex Genome Analysis. Various uses for DNA variations.

At the 4th International Meeting on Single Nucleotide Polymorphism and Complex Genome Analysis (Stockholm, Sweden, 10th-14th October 2001), approximately 100 scientists from more than 20 nations undertook a probing review of latest developments in the field. Despite impressive and still ongoing activities towards SNP discovery and validation, plus efforts towards haplotype exploitation, it was clear that supporting technologies for genotyping are way behind where they need to be. Innate complexity and large variances in aspects of genome function together pose immense challenges that are difficult to surmount in the human situation. In contrast, studies in simpler organisms and population/evolutionary genetics studies are yielding important new insights. Breakthroughs that are being made in understanding the genetic etiology of complex disease tend to involve genes of larger effect or extremely well merited candidates. Linkage studies and proximal phenotypes are being recommended, though the best way forward is still hotly debated. Consequently, many diverse and ambitious projects are underway, from which the data itself will eventually show what is and is not possible.

Genome, Human↗

Genome analysis with gene-indexing databases.

The recent release of the draft sequence and the eventual completion of the human genome present the scientific community with a rich source of data to mine. Yet, these data are content poor in the absence of additional correlative information. Expressed sequence tag (EST) datasets and their associated gene indices have existed for many years, and represent the first attempt at understanding the complexity of the genome. These datasets remain extremely important as information sources and, in particular, as tools for analyzing the completed genomes. Here, we discuss the nature of ESTs and their associated tools and gene-indexing databases. In particular, we will compare three EST gene indices (UNIGENE, Merck Gene Index Version 2.0 and Doubletwist CAT), discuss how these gene indices are applied for both genome analysis and drug discovery, and demonstrate their importance as a complementary dataset to the annotated human genome.

Databases, Factual↗

Comparative genome analysis: selection pressure on the Borrelia vls cassettes is essential for infectivity.

BACKGROUND: At least three species of Borrelia burgdorferi sensu lato (Bbsl) cause tick-borne Lyme disease. Previous work including the genome analysis of B. burgdorferi B31 and B. garinii PBi suggested a highly variable plasmid part. The frequent occurrence of duplicated sequence stretches, the observed plasmid redundancy, as well as the mainly unknown function and variability of plasmid encoded genes rendered the relationships between plasmids within and between species largely unresolvable. RESULTS: To gain further insight into Borreliae genome properties we completed the plasmid sequences of B. garinii PBi, added the genome of a further species, B. afzelii PKo, to our analysis, and compared for both species the genomes of pathogenic and apathogenic strains. The core of all Bbsl genomes consists of the chromosome and two plasmids collinear between all species. We also found additional groups of plasmids, which share large parts of their sequences. This makes it very likely that these plasmids are relatively stable and share common ancestors before the diversification of Borrelia species. The analysis of the differences between B. garinii PBi and B. afzelii PKo genomes of low and high passages revealed that the loss of infectivity is accompanied in both species by a loss of similar genetic material. Whereas B. garinii PBi suffered only from the break-off of a plasmid end, B. afzelii PKo lost more material, probably an entire plasmid. In both cases the vls gene locus encoding for variable surface proteins is affected. CONCLUSION: The complete genome sequences of a B. garinii and a B. afzelii strain facilitate further comparative studies within the genus Borrellia. Our study shows that loss of infectivity can be traced back to only one single event in B. garinii PBi: the loss of the vls cassettes possibly due to error prone gene conversion. Similar albeit extended losses in B. afzelii PKo support the hypothesis that infectivity of Borrelia species depends heavily on the evasion from the host response.

Borrelia↗

Genomic analysis of Grapevine Retrotransposon 1 (Gret 1) in Vitis vinifera.

The complete sequence of the first retrotransposon isolated in Vitis vinifera, Gret 1, was used to design primers that permitted its analysis in the genome of grapevine cultivars. This retroelement was found to be dispersed throughout the genome with sites of repeated insertions. Fluorescent in situ hybridization indicated multiple Gret 1 loci distributed throughout euchromatic portions of chromosomes. REMAP and IRAP proved to be useful as molecular markers in grapevine. Both of these techniques showed polymorphisms between cultivars but not between clones of the same cultivar, indicating differences in Gret 1 distribution between cultivars. The combined cytological and molecular results suggest that Gret 1 may have a role in gene regulation and in explaining the enormous phenotypic variability that exists between cultivars.

Base Sequence↗

Genomic analysis of G protein gamma subunits in human and mouse - the relationship between conserved gene structure and G protein betagamma dimer formation.

Analysis of the genomic sequences, cDNAs and expressed sequence tags (ESTs) in human and mouse for the 12 genes of the gamma subunits of the heterotrimeric G proteins has allowed us to identify the common versus unique elements of the organization and expression of the members of this important gene family. All of the G protein gamma subunit genes are organized around two coding exons, each containing about 100 nucleotides coding for 30-40 amino acids. These two exons each correspond to a functional domain of the protein, which interestingly appears to impose constraints on both the structure of the protein and the structure of the gene. There is large variation in the intron size between these two coding exons, the number and size of 5' and 3' UTRs, and the overall size of the genes. There is general but not absolute conservation in the size and structure of these genes between humans and mice. Alternative splicing and potential differential promoter usage were detected for several Ggamma subunits, indicating possible differential regulation in expression. Only for Ggamma10, however, did we find an alternative coding transcript. This alternative transcript appears to code for a hybrid protein containing a DnaJ domain in place of its Ggamma exon 1 domain, joined to the Ggamma10 second exon domain. The predicted mRNA is expressed in humans, and the protein coded by it is readily translated in vitro. This protein does not form a functional G protein betagamma dimer, but it could generate a chaperone-like protein related to its DNA-J domain. These studies suggest that alternative splicing is not a prominent mechanism for generating G protein subunit diversity from within the human or mouse genomes. Instead, each of the known 12 gamma subunit genes generate transcripts with one prevalent protein.

Alternative Splicing↗

Structure-based active site profiles for genome analysis and functional family subclassification.

In previous work, structure-based functional site descriptors, fuzzy functional forms (FFFs), were developed to recognize structurally conserved active sites in proteins. These descriptors identify members of protein families according to active-site structural similarity, rather than overall sequence or structure similarity. FFFs are defined by a minimal number of highly conserved residues and their three-dimensional arrangement. This approach is advantageous for function assignment across broad families, but is limited when applied to detailed subclassification within these families. In the work described here, we developed a method of three-dimensional, or structure-based, active-site profiling that utilizes FFFs to identify residues located in the spatial environment around the active site. Three-dimensional active-site profiling reveals similarities and differences among active sites across protein families. Using this approach, active-site profiles were constructed from known structures for 193 functional families, and these profiles were verified as distinct and characteristic. To achieve this result, a scoring function was developed that discriminates between true functional sites and those that are geometrically most similar, but do not perform the same function. In a large-scale retrospective analysis of human genome sequences, this profile score was shown to identify specific functional families correctly. The method is effective at recognizing the likely subtype of structurally uncharacterized members of the diverse family of protein kinases, categorizing sequences correctly that were misclassified by global sequence alignment methods. Subfamily information provided by this three-dimensional active-site profiling method yields key information for specific and selective inhibitor design for use in the pharmaceutical industry.

Algorithms↗

Assessing evolutionary relationships among microbes from whole-genome analysis.

The determination and analysis of complete genome sequences have recently enabled many major advances to be made in the area of microbial evolutionary biology. These include the determination of the first genome of a Crenarchaeota, the suggestion that horizontal gene transfer may be the rule rather than the exception, and revelations about how genomes evolve on short timescales.

Archaea↗

RabbitSketch: a high-performance sketching library for genome analysis.

SUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest.

Software↗

Systematic sequencing of the Escherichia coli genome: analysis of the 0-2.4 min region.

A contiguous 111,402-nucleotide sequence corresponding to the 0 to 2.4 min region of the E. coli chromosome was determined as a first step to complete structural analysis of the genome. The resulting sequence was used to predict open reading frames and to search for sequence similarity against the PIR protein database. A number of novel genes were found whose predicted protein sequences showed significant homology with known proteins from various organisms, including several clusters of genes similar to those involved in fatty acid metabolism in bacteria (e.g., betT, baiF) and higher organisms, iron transport (sfuA, B, C) in Serratia marcescens, and symbiotic nitrogen fixation or electron transport (fixA, B, C, X) in Azorhizobium caulinodans. In addition, several genes and IS elements that had been mapped but not sequenced (e.g., leuA, B, C, D) were identified. We estimate that about 90 genes are represented in this region of the chromosome with little spacer.

Bacterial Proteins↗

Diversity and biocatalytic potential of epoxide hydrolases identified by genome analysis.

Epoxide hydrolases play an important role in the biodegradation of organic compounds and are potentially useful in enantioselective biocatalysis. An analysis of various genomic databases revealed that about 20% of sequenced organisms contain one or more putative epoxide hydrolase genes. They were found in all domains of life, and many fungi and actinobacteria contain several putative epoxide hydrolase-encoding genes. Multiple sequence alignments of epoxide hydrolases with other known and putative alpha/beta-hydrolase fold enzymes that possess a nucleophilic aspartate revealed that these enzymes can be classified into eight phylogenetic groups that all contain putative epoxide hydrolases. To determine their catalytic activities, 10 putative bacterial epoxide hydrolase genes and 2 known bacterial epoxide hydrolase genes were cloned and overexpressed in Escherichia coli. The production of active enzyme was strongly improved by fusion to the maltose binding protein (MalE), which prevented inclusion body formation and facilitated protein purification. Eight of the 12 fusion proteins were active toward one or more of the 21 epoxides that were tested, and they converted both terminal and nonterminal epoxides. Four of the new epoxide hydrolases showed an uncommon enantiopreference for meso-epoxides and/or terminal aromatic epoxides, which made them suitable for the production of enantiopure (S,S)-diols and (R)-epoxides. The results show that the expression of epoxide hydrolase genes that are detected by analyses of genomic databases is a useful strategy for obtaining new biocatalysts.

Animals↗

Comparative genomic analysis of plant-associated bacteria.

This review deals with a comparative analysis of seven genome sequences from plant-associated bacteria. These are the genomes of Agrobacterium tumefaciens, Mesorhizobium loti, Sinorhizobium meliloti, Xanthomonas campestris pv campestris, Xanthomonas axonopodis pv citri, Xylella fastidiosa, and Ralstonia solanacearum. Genome structure and the metabolism pathways available highlight the compromise between the genome size and lifestyle. Despite the recognized importance of the type III secretion system in controlling host compatibility, its presence is not universal in all necrogenic pathogens. Hemolysins, hemagglutinins, and some adhesins, previously reported only for mammalian pathogens, are present in most organisms discussed. Different numbers and combinations of cell wall degrading enzymes and genes to overcome the oxidative burst generally induced by the plant host are characterized in these genomes. A total of 19 genes not involved in housekeeping functions were found common to all these bacteria.

Adaptation, Physiological↗

Comparative genomic analysis of the clade B serpin cluster at human chromosome 18q21: amplification within the mouse squamous cell carcinoma antigen gene locus.

The human clade B serpins neutralize serine or cysteine proteinases and reside predominantly within the intracellular compartment. Genomic analysis shows that the 13 human clade B serpins map to either 6p25 (n = 3) or 18q21 (n = 10). Similarly, the mouse clade B serpins map to syntenic loci at 13A3.2 and 1D, respectively. The mouse clade B cluster at 13A3.2 shows a marked expansion in the number of serpin genes (n = 15). The purpose of this study was to determine whether a similar expansion occurred at 1D. Using STS-content mapping, comparative genomic DNA sequence analysis, and cDNA cloning, we found that the mouse clade B cluster at 1D showed nearly complete conservation of gene number, order, and orientation relative to those of 18q21. The only exception was the squamous cell carcinoma antigen (SCCA) locus. The human SCCA locus contains two genes, SERPINB3 (SCCA1) and SERPINB4 (SCCA2), whereas the mouse locus contains four serpins and three pseudogenes. Based on phylogenetic analysis and predicted amino acid sequences, amplification of the mouse SCCA locus occurred after rodents and primates diverged and was associated with some diversification of proteinase inhibitory activity relative to that of humans.

Amino Acid Sequence↗

Longitudinal whole-genome analysis of bluetongue virus identifies conserved serotype-specific genomes and distinct genomic constellations within a Colorado sheep flock (2021-2023).

Bluetongue virus (BTV) is a segmented double-stranded RNA virus of ruminants transmitted by Culicoides spp. biting midges. Although the genome consists of ten segments, classification into serotypes is primarily based on genome segment 2. However, reassortment among genomic segments is a major driver of BTV evolution and diversity. This study used longitudinal whole-genome sequencing to characterize BTV genomes collected from 2021 to 2023 within a single sheep flock in Colorado, where multiple serotypes co-circulate. Whole-genome sequences were generated from fourteen blood samples representing four serotypes: BTV-6, -11, -13, and -17. Longitudinal sampling identified multiple BTV serotypes within individual sheep across consecutive years. Tanglegram analysis comparing segment phylogenies to the segment 2 tree demonstrated incongruent topologies across all genomic segments, suggestive of reassortment or the circulation of distinct genomic constellations. Nucleotide-level comparisons revealed high sequence homology among same-serotype samples from the same year, while the greatest genetic divergence was observed among BTV-17 genomes collected in different years. Additionally, all BTV-13 genomes contained a previously undescribed nonsynonymous substitution in segment 10 predicted to extend the encoded protein by three amino acids. Together, these findings demonstrate that highly conserved BTV genomes and distinct genomic constellations can be detected at the flock level across multiple years. This longitudinal whole-genome approach reveals the genetic complexity of endemic BTV populations, including novel variants and genomic patterns consistent with reassortment that are lost with conventional serotyped-based approaches, highlighting the need to integrate whole-genome characterization into endemic BTV monitoring programs.

Animals↗

Applications of DNA tiling arrays for whole-genome analysis.

DNA microarrays are a well-established technology for measuring gene expression levels. Microarrays designed for this purpose use relatively few probes for each gene and are biased toward known and predicted gene structures. Recently, high-density oligonucleotide-based whole-genome microarrays have emerged as a preferred platform for genomic analysis beyond simple gene expression profiling. Potential uses for such whole-genome arrays include empirical annotation of the transcriptome, chromatin-immunoprecipitation-chip studies, analysis of alternative splicing, characterization of the methylome (the methylation state of the genome), polymorphism discovery and genotyping, comparative genome hybridization, and genome resequencing. Here we review different whole-genome microarray designs and applications of this technology to obtain a wide variety of genomic scale information.

Animals↗