Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Moving toward whole-genome analysis: a technology perspective.

PURPOSE: New, highly efficient technologies used in genomic analysis are described, and their implications for health care are discussed. SUMMARY: The availability of the human genome sequence, in confluence with the ability to affordably package it for analysis, is opening new frontiers in biomedical research. On the horizon, personalized medicine--driven by molecular characterization of disease, genetic analysis of the patient, and information technologies designed to enable health care professionals to leverage these tools--promises to fundamentally transform health care. New genetics technologies, such as high-density microarrays, will fuel this research by providing researchers with the ability to comprehensively access the human genome in all its complexity. Some of the most promising areas for application of genetic information are those where society's current needs are greatest: complex, common disorders, such as cancer and cardiovascular disease; drug interactions; inherited genetic disorders that afflict children; and late-onset conditions for which no cure currently exists. The barriers to using genetic information widely in health care are in many cases not technological or economic, but social and political. CONCLUSION: New technology enables efficient, large-scale analysis of the whole genome, genetic variations, and gene expression. Genomic analysis has profound clinical, economic, and social implications for health care.

Biomedical Research↗

GNARE: automated system for high-throughput genome analysis with grid computational backend.

Recent progress in genomics and experimental biology has brought exponential growth of the biological information available for computational analysis in public genomics databases. However, applying the potentially enormous scientific value of this information to the understanding of biological systems requires computing and data storage technology of an unprecedented scale. The Grid, with its aggregated and distributed computational and storage infrastructure, offers an ideal platform for high-throughput bioinformatics analysis. To leverage this we have developed the Genome Analysis Research Environment (GNARE)--a scalable computational system for the high-throughput analysis of genomes, which provides an integrated database and computational backend for data-driven bioinformatics applications. GNARE efficiently automates the major steps of genome analysis including acquisition of data from multiple genomic databases; data analysis by a diverse set of bioinformatics tools; and storage of results and annotations. High-throughput computations in GNARE are performed using distributed heterogeneous Grid computing resources such as Grid2003, TeraGrid, and the DOE Science Grid. Multi-step genome analysis workflows involving massive data processing, the use of application-specific tools and algorithms and updating of an integrated database to provide interactive web access to results are all expressed and controlled by a "virtual data" model which transparently maps computational workflows to distributed Grid resources. This paper describes how Grid technologies such as Globus, Condor, and the Gryphyn Virtual Data System were applied in the development of GNARE. It focuses on our approach to Grid resource allocation and to the use of GNARE as a computational framework for the development of bioinformatics applications.

Computational Biology↗

Comparative genomic analysis of three strains of Ehrlichia ruminantium reveals an active process of genome size plasticity.

Ehrlichia ruminantium is the causative agent of heartwater, a major tick-borne disease of livestock in Africa that has been introduced in the Caribbean and is threatening to emerge and spread on the American mainland. We sequenced the complete genomes of two strains of E. ruminantium of differing phenotypes, strains Gardel (Erga; 1,499,920 bp), from the island of Guadeloupe, and Welgevonden (Erwe; 1,512,977 bp), originating in South Africa and maintained in Guadeloupe in a different cell environment. Comparative genomic analysis of these two strains was performed with the recently published parent strain of Erwe (Erwo) and other Rickettsiales (Anaplasma, Wolbachia, and Rickettsia spp.). Gene order is highly conserved between the E. ruminantium strains and with A. marginale. In contrast, there is very little conservation of gene order with members of the Rickettsiaceae. However, gene order may be locally conserved, as illustrated by the tuf operons. Eighteen truncated protein-encoding sequences (CDSs) differentiate Erga from Erwe/Erwo, whereas four other truncated CDSs differentiate Erwe from Erwo. Moreover, E. ruminantium displays the lowest coding ratio observed among bacteria due to unusually long intergenic regions. This is related to an active process of genome expansion/contraction targeted at tandem repeats in noncoding regions and based on the addition or removal of ca. 150-bp tandem units. This process seems to be specific to E. ruminantium and is not observed in the other Rickettsiales.

Conserved Sequence↗

Identification of Leptospira interrogans strains by monoclonal antibodies and genomic analysis.

A recombinant probe derived from a genomic library of serovar hardjo strain Hardjoprajitno, and a panel of serovar specific Monoclonal Antibodies (MAbs) were used for the characterization of 31 Leptospira isolates from cattle and swine. The two methods performed equally well in serovar identification except for the distinction of the genotypes hardjoprajitno and hardjobovis within serovar hardjo which could only be obtained by genomic analysis. The combination of immunological and genetic information was also useful to evaluate the degree of variability of Leptospira strains. The quality of the patterns and the sensitivity provided by a digoxigenin labelled probe were comparable to those obtained with a radioactive reagent.

Animals↗

Comparative genome analysis of Campylobacter jejuni using whole genome DNA microarrays.

Whole genome DNA microarrays were constructed and used to investigate genomic diversity in 18 Campylobacter jejuni strains from diverse sources. New algorithms were developed that dynamically determine the boundary between the conserved and variable genes. Seven hypervariable plasticity regions (PR) were identified in the genome (PR1 to PR7) containing 136 genes (50%) of the variable gene pool. When comparisons were made with the sequenced strain NCTC11168, the number of absent or divergent genes ranged from 2.6% (40 genes) to 10.2% (163) and in total 16.3% (269) of the genes were variable. PR1 contains genes important in the utilisation of alternative electron acceptors for respiration and may confer a selective advantage to strains in restricted oxygen environments. PR2, 3 and 7 contain many outer membrane and periplasmic proteins and hypothetical proteins of unknown function that might be linked to phenotypic variation and adaptation to different ecological niches. PR4, 5 and 6 contain genes involved in the production and modification of antigenic surface structures.

Algorithms↗

Incorporating Epidemiological Data into the Genomic Analysis of Partially Sampled Infectious Disease Outbreaks.

Pathogen genomic data are increasingly being used to investigate transmission dynamics in infectious disease outbreaks. Combining genomic data with epidemiological data should substantially increase our understanding of outbreaks, but this is highly challenging when the outbreak under study is only partially sampled, so that both genomic and epidemiological data are missing for intermediate links in the transmission chains. Here, we present a new dynamic programming algorithm to perform this task efficiently. We implement this methodology into the well-established TransPhylo framework to reconstruct partially sampled outbreaks using a combination of genomic and epidemiological data. We use simulated datasets to show that including epidemiological data can improve the accuracy of the inferred transmission links compared with inference based on genomic data only. This also allows us to estimate parameters specific to the epidemiological data (such as transmission rates between particular groups), which would otherwise not be possible. We then apply these methods to two real-world examples. First, we use genomic data from an outbreak of tuberculosis in Argentina, for which data was also available on the HIV status of sampled individuals, in order to investigate the role of HIV coinfection in the spread of this tuberculosis outbreak. Second, we use genomic and geographical data from the 2003 epidemic of avian influenza H7N7 in the Netherlands to reconstruct its spatial epidemiology. In both cases, we show that incorporating epidemiological data into the genomic analysis allows us to investigate the role of epidemiological properties in the spread of infectious diseases.

Humans↗

Molecular characterization of a t(9;12)(p21;q13) balanced chromosome translocation in combination with integrative genomics analysis identifies C9orf14 as a candidate tumor-suppressor.

A large number of nevi (LNN) is a high risk phenotypic trait for developing cutaneous malignant melanoma (CMM). In this study, the breakpoints of a t(9;12)(p21;q13) balanced chromosome translocation were finely mapped in a family with LNN and CMM. Molecular characterization of the 9p21 breakpoint identified a novel gene C9orf14 expressed in melanocytes disrupted by the translocation. Integrative analysis of functional genomics data was applied to determine the role of C9orf14 in CMM development. An analysis of genome-wide DNA copy number alterations in melanoma tumors revealed the loss of the C9orf14 locus, located proximal to CDKN2A, in approximately one-fourth of tumors. Analysis of gene expression data in cancer cell lines and melanoma tumors suggests a loss of C9orf14 expression in melanoma tumorigenesis. Taken together, our results indicate that C9orf14 is a candidate tumor-suppressor for nevus development and late stage melanoma at 9p21, a region frequently deleted in different types of human cancers.

Chromosomes, Human, Pair 12↗

Microbial genome analysis: insights into virulence, host adaptation and evolution.

Genome analysis of microbial pathogens has provided unique insights into their virulence, host adaptation and evolution. Common themes have emerged, including lateral gene transfer among enteric pathogens, genome decay among obligate intracellular pathogens and antigenic variation among mucosal pathogens. The advent of post-genomic approaches and the sequencing of the human genome will enable scientists to investigate the complex and dynamic interplay between host and pathogen. This wealth of information will catalyse the development of new intervention strategies to reduce the burden of microbial-related disease.

Adaptation, Physiological↗

Pitfalls of bacterial pan-genome analysis approaches: a case study of Mycobacterium tuberculosis and two less clonal bacterial species.

SUMMARY: Pan-genome analysis is a fundamental tool for studying bacterial genome evolution; however, the variety in methods used to define and measure the pan-genome poses challenges to the interpretation and reliability of results. Using Mycobacterium tuberculosis, a clonally evolving bacterium with a small accessory genome, as a model system, we systematically evaluated sources of variability in pan-genome estimates. Our analysis revealed that differences in assembly type (short-read versus hybrid), annotation pipeline, and pan-genome software, significantly impact predictions of core and accessory genome size. Extending our analysis to two additional bacterial species, Escherichia coli and Staphylococcus aureus, we observed consistent tool-dependent biases but species-specific patterns in pan-genome variability. Our findings highlight the importance of integrating nucleotide- and protein-level analyses to improve the reliability and reproducibility of pan-genome studies across diverse bacterial populations. AVAILABILITY AND IMPLEMENTATION: Panqc is freely available under an MIT license at https://github.com/maxgmarin/panqc.

Genome, Bacterial↗

Use of genomic analysis of varicella-zoster virus to investigate suspected varicella-zoster transmission within a renal unit.

BACKGROUND: The source of hospital-acquired chickenpox infection may be presumed from a known exposure, but has not been previously proven using genomic analysis. OBJECTIVE: Investigation of suspected VZV transmission was done using single nucleotide polymorphism genomic analysis. STUDY DESIGN: Comparison was made of viral isolates from two patients with chickenpox on the same ward who were not known to have had direct contact. RESULTS: An identical genotype in the variable R1 region of the VZV was isolated from the two patients. CONCLUSION: Inapparent hospital-acquired transmission was the most likely route of infection.

Acyclovir↗

Thermus thermophilus genome analysis: benefits and implications.

The genome sequence analysis of Thermus thermophilus HB27, a microorganism with high biotechnological potential, has recently been published. In that report, the chromosomal and the megaplasmid sequence were compared to those of other organisms and discussed on the basis of their physiological and metabolic features. Out of the 2,218 putative genes identified through the large genome sequencing project, a significant number has potential interest for biotechnology. The present communication will discuss the accumulating information on molecules participating in fundamental biological processes or having potential biotechnological importance.

Editorial↗

Multilocus markers for mouse genome analysis: PCR amplification based on single primers of arbitrary nucleotide sequence.

Polymerase chain reaction (PCR) based on single primers of arbitrary nucleotide sequence provides a powerful marker system for genome analysis because each primer amplifies multiple products, and cloning, sequencing, and hybridization are not required. We have evaluated this typing system for the mouse by identifying optimal PCR conditions; characterizing effects of GC content, primer length, and multiplexed primers; demonstrating considerable variation among a panel of inbred strains; and establishing linkage for several products. Mg2+, primer, template, and annealing conditions were identified that optimized the number and resolution of amplified products. Primers with 40% GC content failed to amplify products readily, primers with 50% GC content resulted in reasonable amplification, and primers with 60% GC content gave the largest number of well-resolved products. Longer primers did not necessarily amplify more products than shorter primers of the same proportional GC content. Multiplexed primers yielded more products than either primer alone and usually revealed novel variants. A strain survey showed that most strains could be readily distinguished with a modest number of primers. Finally, linkage for seven products was established on five chromosomes. These characteristics establish single primer PCR as a powerful method for mouse genome analysis.

Animals↗

A comparative genomic analysis of ESTs from Ustilago maydis.

A large-scale comparative genomic analysis of unisequence sets obtained from an Ustilago maydis EST collection was performed against publicly available EST and genomic sequence datasets from 21 species. We annotated 70% of the collection based on similarity to known sequences and recognized protein signatures. Distinct grouping of the ESTs, defined by the presence or absence of similar sequences in the species examined, allowed the identification of U. maydis sequences present only (1) in fungal species, (2) in plants but not animals, (3) in animals but not plants, or (4) in all three eukaryotic lineages assessed. We also identified 215 U. maydis genes that are found in the ascomycete but not in the basidiomycete genome sequences searched. Candidate genes were identified for further functional characterization. These include 167 basidiomycete-specific sequences, 58 fungal pathogen-specific sequences (including 37 basidiomycete pathogen-specific sequences), and 18 plant pathogen-specific sequences, as well as two sequences present only in other plant pathogen and plant species.

Databases, Nucleic Acid↗

A hybrid gene team model and its application to genome analysis.

It is well-known that functionally related genes occur in a physically clustered form, especially operons in bacteria. By leveraging on this fact, there has recently been an interesting problem formulation known as gene team model, which searches for a set of genes that co-occur in a pair of closely related genomes. However, many gene teams, even experimentally verified operons, frequently scatter within other genomes. Thus, the gene team model should be refined to reflect this observation. In this paper, we generalized the gene team model, that looks for gene clusters in a physically clustered form, to multiple genome cases with relaxed constraints. We propose a novel hybrid pattern model that combines the set and the sequential pattern models. Our model searches for gene clusters with and/or without physical proximity constraint. This model is implemented and tested with 97 genomes (120 replicons). The result was analyzed to show the usefulness of our model. We also compared the result from our hybrid model to those from the traditional gene team model. We also show that predicted gene teams can be used for various genome analysis: operon prediction, phylogenetic analysis of organisms, contextual sequence analysis and genome annotation. Our program is fast enough to provide a service on the web at http://platcom.informatics.indiana.edu/platcom/. Users can select any combination of 97 genomes to predict gene teams.

Algorithms↗

Using bioinformatics and genome analysis for new therapeutic interventions.

The genome era provides two sources of knowledge to investigators whose goal is to discover new cancer therapies: first, information on the 20,000 to 40,000 genes that comprise the human genome, the proteins they encode, and the variation in these genes and proteins in human populations that place individuals at risk or that occur in disease; second, genome-wide analysis of cancer cells and tissues leads to the identification of new drug targets and the design of new therapeutic interventions. Using genome resources requires the storage and analysis of large amounts of diverse information on genetic variation, gene and protein functions, and interactions in regulatory processes and biochemical pathways. Cancer bioinformatics deals with organizing and analyzing the data so that important trends and patterns can be identified. Specific gene and protein targets on which cancer cells depend can be identified. Therapeutic agents directed against these targets can then be developed and evaluated. Finally, molecular and genetic variation within a population may become the basis of individualized treatment.

Computational Biology↗

Genome analysis of adenovirus 4 isolated over a six year period.

Genome analysis was carried out on 74 adenovirus 4 (Ad4) isolates from patients in Manchester between 1984 and 1989. Most of the isolates were associated with conjunctivitis. Of the 74 isolates studied, 51 were Ad4a and 10 were Ad4p (the prototype strain). The remaining isolates consisted of two new genome types we have designated Ad4a2 (10 isolates) and Ad4a3 (3 isolates). Most of the genome types co-circulated during the period of study. The Bst E II and Xho I restriction maps of the new variants are presented and compared with those of Ad4p. We are unable to associate genome types with particular clinical presentations.

Adenovirus Infections, Human↗

Assessing the Safety and Probiotic Potential of Bifidobacterium longum subsp. infantis BI45: A Comprehensive Study from Genomic Analysis to Randomized Controlled Clinical Trial in Healthy Adults.

While probiotics are increasingly consumed for health benefits, comprehensive safety assessments, particularly for novel strains, are imperative. This study aimed to conduct a holistic safety and efficacy assessment of Bifidobacterium longum subsp. infantis (B. infantis) BI45, spanning genomic analysis, in vitro tests, In vivo toxicity test, and a clinical trial. The safety of B. infantis BI45 was evaluated through: (1) whole-genome sequencing for antibiotic resistance and virulence genes; (2) in vitro phenotyping (hemolysis, cytotoxicity, gastrointestinal tolerance and antibiotic susceptibility); (3) an acute oral toxicity study in mice; and (4) a randomized, double-blind, placebo-controlled clinical trial. Forty-eight healthy adults were recruited and randomly assigned to receive either B. infantis BI45 or a placebo (n&#x2009;=&#x2009;24/group) for 8 weeks. Hematological, biochemical, immunological, and gut microbiota parameters were assessed. Genomic analysis identified no transferable antibiotic resistance or virulence genes. In vitro assays confirmed the absence of hemolytic and cytotoxic activity, alongside high gastrointestinal tolerance. Antibiotic susceptibility testing showed that B. infantis BI45 is sensitive to a range of antibiotics. No adverse effects were observed in the murine toxicity study at 2&#x2009;&#xd7;&#x2009;10&#xb9;&#x2070; CFU/kg. Importantly, the clinical intervention revealed no adverse events or significant alterations in hematological, hepatic, or renal function markers in the B. infantis BI45 group, demonstrating an excellent safety profile. Furthermore, B. infantis BI45 supplementation significantly increased serum levels of immunomodulatory markers Immunoglobulin A (IgA) and antimicrobial peptide LL-37 compared to the placebo (p&#x2009;<&#x2009;0.05) and modulated the gut microbiota by enriching beneficial short-chain fatty acid producers. The multi-tiered evidence demonstrates that B. infantis BI45 is a safe probiotic strain that does not induce adverse reactions in healthy adults. Its consumption positively modulates host immunity and the gut microbiota.Trial Registration Number: NCT06863415 (ClinicalTrials.gov).

Adult↗