Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Genomic analysis of facioscapulohumeral muscular dystrophy.

The genomic basis of facioscapulohumeral muscular dystrophy (FSHD) is of considerable interest because of the unique nature of the molecular mutation, which is a deletion within a large, complex DNA tandem array (D4Z4). This repeat maps within 30 kb of the 4q telomere. Although D4Z4 repeat units each contain an open reading frame that could encode a homeodomain protein, there is no evidence that the repeat is transcribed, and the underlying disease mechanism probably involves a position effect. A recent study has identified a protein complex bound to D4Z4 that contains YY1 and HMGB2, implicating a role for D4Z4 as a repressor. The 4q telomere has two variants, 4qA and 4qB. Although these alleles are present at almost equal frequencies in the general population, FSHD is associated only with the 4qA allele and never with 4qB. This suggests a functional difference between the telomere variants, either in predisposition to deletions within D4Z4 or in the pathological consequence of the deletion. Comparative mapping studies of the FSHD region in primates, mouse and Fugu rubripes have given insights into the evolutionary history of the D4Z4 repeat and of 4qter, although as yet they have not provided any solutions to the FSHD puzzle.

Animals↗

Automated bacterial genome analysis and annotation.

More than 300 bacterial genome sequences are publicly available, and many more are scheduled to be completed and released in the near future. Converting this raw sequence information into a better understanding of the biology of bacteria involves the identification and annotation of genes, proteins and pathways. This processing is typically done using sequence annotation pipelines comprised of a variety of software modules and, in some cases, human experts. The reference databases, computational methods and knowledge that form the basis of these pipelines are constantly evolving, and thus there is a need to reprocess genome annotations on a regular basis. The combined challenge of revising existing annotations and extracting useful information from the flood of new genome sequences will necessitate more reliance on completely automated systems.

Bacterial Proteins↗

Regional genomic analysis of lineage distribution and transferable multidrug resistance among chicken-associated Salmonella Kentucky isolates in China.

Salmonella enterica serovar Kentucky is an important multidrug-resistant foodborne pathogen in the poultry meat supply chain. Although recent broader genomic studies have elucidated the population structure and epidemiological significance of major lineages in China (e.g., ST198 and ST314), the regional dynamics within local poultry supply chains remain insufficiently characterized. In this study, 31 chicken meat-derived isolates from Shanghai and 39 publicly available genomes from China were analyzed using antimicrobial susceptibility testing, whole-genome sequencing, phylogenetic analysis, conjugation experiments, and complete sequencing of representative plasmids. This enabled a systematic characterization of the molecular epidemiological features of the population and the mechanisms underlying resistance dissemination. Population genomic analysis revealed a lineage composition markedly different from the global epidemiological pattern: ST314 was the predominant sequence type among the Shanghai chicken-derived isolates (74.2%), whereas the internationally recognized high-risk clone ST198 accounted for only 25.8% of the local isolates. However, risk stratification analysis indicated that although ST198 was detected less frequently, it carried a significantly greater burden of acquired resistance genes and therefore represented a higher-risk resistant lineage. Functional and structural validation further elucidated the molecular basis of resistance dissemination within this high-risk lineage. Conjugation experiments confirmed the co-transfer of a multidrug resistance module carrying blaTEM-1 and blaCTX-M-267 to the recipient strain Escherichia coli J53. Complete plasmid analysis revealed that these two β-lactam resistance genes were co-localized on a 242-kb transferable plasmid flanked by Tn1331, Tn3, and multiple transposase-associated elements, thereby providing a structural basis for their horizontal transfer. This study provides important molecular epidemiological evidence for lineage-specific surveillance and risk-stratified control of resistant Salmonella in the poultry meat supply chain and further underscores the need for continuous monitoring of mobile genetic elements within a One Health framework.

Animals↗

Functional genomics analysis of Singapore grouper iridovirus: complete sequence determination and proteomic analysis.

Here we report the complete genome sequence of Singapore grouper iridovirus (SGIV). Sequencing of the random shotgun and restriction endonuclease genomic libraries showed that the entire SGIV genome consists of 140,131 nucleotide bp. One hundred sixty-two open reading frames (ORFs) from the sense and antisense DNA strands, coding for lengths varying from 41 to 1,268 amino acids, were identified. Computer-assisted analyses of the deduced amino acid sequences revealed that 77 of the ORFs exhibited homologies to known virus genes, 23 of which matched functional iridovirus proteins. Forty-two putative conserved domains or signatures were detected in the National Center for Biotechnology Information CD-Search database and PROSITE database. An assortment of enzyme activities involved in DNA replication, transcription, nucleotide metabolism, cell signaling, etc., were identified. Viruses were cultured on a cell line derived from the embryonated egg of the grouper Epinephelus tauvina, isolated, and purified by sucrose gradient ultracentrifugation. The protein extract from the purified virions was analyzed by polyacrylamide gel electrophoresis followed by in-gel digestion of protein bands. Matrix-assisted laser desorption ionization-time of flight mass spectrometry and database searching led to identification of 26 proteins. Twenty of these represented novel or previously unidentified genes, which were further confirmed by reverse transcription-PCR (RT-PCR) and DNA sequencing of their respective RT-PCR products.

Animals↗

Systematic sequencing of the Escherichia coli genome: analysis of the 2.4-4.1 min (110,917-193,643 bp) region.

The complete sequence analysis of the E. coli genome was initiated as a collaborative study in Japan. Following the initial analysis of the 0-2.4 min region (Yura, T. et al. (1992) Nucleic Acids Res. 20, 3305-3308), a contiguous sequence of 82,727 bp corresponding to the 2.4-4.1 min region (110,917-193,643 bp as counted from 0 min) was determined. The resulting sequence was found to contain at least 33 known genes and 24 putative genes predicted from protein sequence homology.

Bacterial Proteins↗

Translation in Bacillus subtilis: roles and trends of initiation and termination, insights from a genome analysis.

We analysed the Bacillus subtilis protein coding sequences termini, and compared it to other genomes. The analysis focused on signals, com-positional biases of nucleotides, oligonucleotides, codons and amino acids and mRNA secondary structure. AUG is the preferred start codon in all genomes, independent of their G+C content, and seems to induce less stable mRNA structures. However, it is not conserved between homologous genes neither is it preferred in highly expressed genes. In B.subtilis the ribosome binding site is very strong. We found that downstream boxes do not seem to exist either in Escherichia coli or in B.subtilis. UAA stop codon usage is correlated with the G+C content and is strongly selected in highly expressed genes. We found less stable mRNA structures at both termini, which we related to mRNA-ribosome and mRNA-release-factor interactions. This pattern seems to impose a peculiar A-rich nucleotide and codon usage bias in these regions. Finally the analysis of all proteins from B.subtilis revealed a similar amino acid bias near both termini of proteins consisting of over-representation of hydrophilic residues. This bias near the stop codon is partially release-factor specific.

Algorithms↗

A genomic analysis of rat proteases and protease inhibitors.

Proteases perform important roles in multiple biological and pathological processes. The availability of the rat genome sequence has facilitated the analysis of the complete protease repertoire or degradome of this model organism. The rat degradome consists of at least 626 proteases and homologs, which are distributed into 24 aspartic, 160 cysteine, 192 metallo, 221 serine, and 29 threonine proteases. This distribution is similar to that of the mouse degradome but is more complex than that of the human degradome composed of 561 proteases and homologs. This increased complexity of rat proteases mainly derives from the expansion of several families, including placental cathepsins, testases, kallikreins, and hematopoietic serine proteases, involved in reproductive or immunological functions. These protease families have also evolved differently in rat and mouse and may contribute to explain some functional differences between these closely related species. Likewise, genomic analysis of rat protease inhibitors has shown some differences with mouse protease inhibitors and the expansion of families of cysteine and serine protease inhibitors in rodents with respect to human. These comparative analyses may provide new views on the functional diversity of proteases and inhibitors and contribute to the development of innovative strategies for treating proteolysis diseases.

Animals↗

Relational genome analysis using reference libraries and hybridisation fingerprinting.

The genomes of eukaryotic organisms are studied by an integrated approach based on hybridisation techniques. For this purpose, a reference library system has been set up, with a wide range of clone libraries made accessible to probe hybridisation as high density filter grids. Many different library types made from a variety of organisms can thus be analysed in a highly parallel process; hence, the amount of work per individual clone is minimised. In addition, information produced on one analysis level instantly assists in the characterisation process on another level. Genetic, physical and transcriptional mapping information and partial sequencing data are obtained for the individual library clones and are cross-referenced toward a comprehensive molecular understanding of genome structure and organisation, of encoded functions and their regulation. The order of genomic clones is established by hybridisation fingerprinting procedures. On these physical maps, the location of transcripts is determined. Complementary, partial sequence information is produced from corresponding cDNAs by hybridising short oligonucleotides, which will lead to the identification of regions of sequence conservation and the constitution of a gene inventory. The hybridisation analysis of the cDNA clones, and the genomic clones as well, could potentially be expanded toward a determination of (nearly) the complete sequence. The accumulated data set will provide the means to direct large-scale sequencing of the DNA, or might even make the sequence analysis of large genomic regions a redundant undertaking due to the already collected information.

Animals↗

The Arabidopsis genome sequence as a tool for genome analysis in Brassicaceae. A comparison of the Arabidopsis and Capsella rubella genomes.

The annotated Arabidopsis genome sequence was exploited as a tool for carrying out comparative analyses of the Arabidopsis and Capsella rubella genomes. Comparison of a set of random, short C. rubella sequences with the corresponding sequences in Arabidopsis revealed that aligned protein-coding exon sequences differ from aligned intron or intergenic sequences in respect to the degree of sequence identity and the frequency of small insertions/deletions. Molecular-mapped markers and expressed sequence tags derived from Arabidopsis were used for genetic mapping in a population derived from an interspecific cross between Capsella grandiflora and C. rubella. The resulting eight Capsella linkage groups were compared to the sequence maps of the five Arabidopsis chromosomes. Fourteen colinear segments spanning approximately 85% of the Arabidopsis chromosome sequence maps and 92% of the Capsella genetic linkage map were detected. Several fusions and fissions of chromosomal segments as well as large inversions account for the observed arrangement of the 14 colinear blocks in the analyzed genomes. In addition, evidence for small-scale deviations from genome colinearity was found. Colinearity between the Arabidopsis and Capsella genomes is more pronounced than has been previously reported for comparisons between Arabidopsis and different Brassica species.

Arabidopsis↗

Sex genes for genomic analysis in human brain: internal controls for comparison of probe level data extraction.

BACKGROUND: Genomic studies of complex tissues pose unique analytical challenges for assessment of data quality, performance of statistical methods used for data extraction, and detection of differentially expressed genes. Ideally, to assess the accuracy of gene expression analysis methods, one needs a set of genes which are known to be differentially expressed in the samples and which can be used as a "gold standard". We introduce the idea of using sex-chromosome genes as an alternative to spiked-in control genes or simulations for assessment of microarray data and analysis methods. RESULTS: Expression of sex-chromosome genes were used as true internal biological controls to compare alternate probe-level data extraction algorithms (Microarray Suite 5.0 [MAS5.0], Model Based Expression Index [MBEI] and Robust Multi-array Average [RMA]), to assess microarray data quality and to establish some statistical guidelines for analyzing large-scale gene expression. These approaches were implemented on a large new dataset of human brain samples. RMA-generated gene expression values were markedly less variable and more reliable than MAS5.0 and MBEI-derived values. A statistical technique controlling the false discovery rate was applied to adjust for multiple testing, as an alternative to the Bonferroni method, and showed no evidence of false negative results. Fourteen probesets, representing nine Y- and two X-chromosome linked genes, displayed significant sex differences in brain prefrontal cortex gene expression. CONCLUSION: In this study, we have demonstrated the use of sex genes as true biological internal controls for genomic analysis of complex tissues, and suggested analytical guidelines for testing alternate oligonucleotide microarray data extraction protocols and for adjusting multiple statistical analysis of differentially expressed genes. Our results also provided evidence for sex differences in gene expression in the brain prefrontal cortex, supporting the notion of a putative direct role of sex-chromosome genes in differentiation and maintenance of sexual dimorphism of the central nervous system. Importantly, these analytical approaches are applicable to all microarray studies that include male and female human or animal subjects.

Algorithms↗

Genome analysis of Theileria parva.

Recent technological developments in the field of genome analyses have advanced our knowledge of the structures of prokaryotic and eukoryotic genomes. Examples of these range from small bacterial genomes, such as Escherichia coli, to the more complex genomes of Caenorhabditis elegans, Drosophila, humans and mouse. Here, Subhash Morzona and John Young review developments in mapping the genome of on economically important protozoan parasite o f cattle, Theileria parva. This map provides a framework for more detailed analysis of the genome structure o f this organism. The methodologies developed in constructing the map also have application to the mapping of other protozoan genomes.

Journal Article↗

Plant genome analysis: the state of the art.

Plants are the basis for the survival of all "higher" organisms on Earth. Development of molecular genetics tools has allowed analysis of the structure, evolution, and function of whole plant genomes, rather than individual genes. DNA-based markers were instrumental in constructing detailed genetic maps of model plants and all major crop species. These molecular maps were the basis of physical maps and the first plant whole genome sequences. Comparative analysis based on genetic, cytogenetic, and physical maps and DNA sequence information provided new insights into the evolution of plant nuclear and organellar genomes. Mapping factors controlling Mendelian and quantitative traits made possible the cloning and functional characterization of novel genes, which function in plant development, adaptation to biotic and abiotic stress, or in the formation of other agronomic characters. The parallel analysis of all transcripts, proteins, and metabolites present in plant cells or tissues has generated information that may lead to a better integrated understanding of genome function. Postfunctional analysis of natural variation of gene function and its effects on phenotype is envisaged to provide new diagnostic and therapeutic molecular tools for applications in plant breeding, adaptation, and ecology.

DNA, Plant↗

Trends in genomic analysis of the cardiovascular system.

OBJECTIVE: To evaluate the opportunities afforded cardiovascular medicine by the comprehensive and integrative approaches of genomics in cellular physiology. We present a meta-analysis of recently reported results obtained by means of high-throughput technologies (complementary DNA and oligonucleotide arrays, serial analysis of gene expression [SAGE]), as well as more traditional molecular biology approaches (real-time polymerase chain reaction, differential display, and others). DATA SOURCES: Newly published articles identified on PubMed and additional data provided by authors on-line (where available). CONCLUSIONS: The impact of genomic analysis on cardiovascular research is already visible. New genes of cardiovascular interest have been discovered, while a number of known genes have been found to be changed in unexpected contexts. The patterns in the variation of expression of many genes correlate well with the models currently used to explain the pathogenesis of cardiovascular diseases. Much more work has yet to be done, however, for the full exploitation of the immense informative potential still dormant in the genomic technologies.

Cardiovascular Diseases↗

Expression analysis, genomic structure, and mapping to 7q31 of the human sperm adhesion molecule gene SPAM1.

During the course of systematic sequence tag analysis of clones isolated from an adult testis cDNA library, clones 296 and 576 were found to detect 71-74% sequence identity to the guinea pig sperm surface protein PH-20. This surface protein is involved in sperm-egg adhesion in the guinea pig. Nucleotide sequence for 1919 bp of human DNA from a series of overlapping cDNA clones isolated from a testis cDNA library confirmed the sequence identity within a 1527-bp open reading frame to be 71-74% to the guinea pig gene and the similarity to be 60% for the predicted protein of 509 amino acids. Southern blot analysis of human genomic DNA and DNA from somatic cell hybrids indicates that the gene (SPAM1) is unique and does not form part of a larger family and that it maps to chromosome 7. Fluorescence in situ hybridization with yeast artificial chromosome (YAC) clones isolated from the CEPH megaYAC library has refined this localization to 7q31. PCR analysis of genomic DNA and YAC clone DNA has shown that the 1919 bp of the gene that has been cloned covers approximately 11 kb of genomic DNA and is encoded by at least 4 exons. Northern analysis of poly(A)+ mRNA from a range of 16 human tissues has demonstrated that expression of the gene as a single 2.4-kb transcript is strictly limited to the testis.

Adult↗

Genomic analysis of prostate carcinoma specimens obtained via ultrasound-guided needle biopsy may be of use in preoperative decision-making.

BACKGROUND: The widespread use of prostate-specific antigen (PSA) testing to screen for prostate carcinoma has led to significant overdiagnosis, due to the frequent detection of indolent malignancies on PSA screening. The detection of abnormal PSA levels typically is followed by ultrasound-guided needle biopsy. Therefore, in an effort to identify genetic markers that augment the information provided by standard histopathologic classification, the authors tested the feasibility of using these minute biopsy samples for genomic profiling via chromosome banding analysis and comparative genomic hybridization (CGH). METHODS: Ultrasound-guided needle biopsy specimens obtained preoperatively from 35 patients with prostate carcinoma were analyzed via chromosome banding analysis (after short-term culturing) and CGH. The findings of these analyses then were analyzed for potential correlations with clinicopathologic parameters. RESULTS: Chromosome banding analysis and CGH were possible in 34 and 33 of the 35 study specimens, respectively. Combined analysis revealed aberrations in 69% of all samples investigated. Copy number losses occurred most commonly at 8p (58% of all abnormal specimens), 16q (42%), and 13q (37%), whereas the only gains detected in more than 1 specimen were those that occurred at 8q (37%). Genomic imbalances and losses at 16q were significantly associated with more poorly differentiated subtypes of prostate carcinoma (P = 0.048 and P = 0.019, respectively), whereas gains at 8q and losses at 16q were significantly correlated with clinically advanced disease (P = 0.048 for the finding of a gain at 8q together with a loss at 16q; P = 0.01 for the finding of either aberration alone). CONCLUSIONS: The authors conclude that genomic analysis of suspected prostate carcinoma specimens obtained via ultrasound-guided needle biopsy is feasible. Thus, it may be possible to use genetic markers to obtain diagnostic and/or prognostic information that is useful in the making of preoperative decisions regarding prostate carcinoma management.

Aged↗

Comparative genomic analysis reveals a novel mitochondrial isoform of human rTS protein and unusual phylogenetic distribution of the rTS gene.

BACKGROUND: The rTS gene (ENOSF1), first identified in Homo sapiens as a gene complementary to the thymidylate synthase (TYMS) mRNA, is known to encode two protein isoforms, rTSalpha and rTSbeta. The rTSbeta isoform appears to be an enzyme responsible for the synthesis of signaling molecules involved in the down-regulation of thymidylate synthase, but the exact cellular functions of rTS genes are largely unknown. RESULTS: Through comparative genomic sequence analysis, we predicted the existence of a novel protein isoform, rTS, which has a 27 residue longer N-terminus by virtue of utilizing an alternative start codon located upstream of the start codon in rTSbeta. We observed that a similar extended N-terminus could be predicted in all rTS genes for which genomic sequences are available and the extended regions are conserved from bacteria to human. Therefore, we reasoned that the protein with the extended N-terminus might represent an ancestral form of the rTS protein. Sequence analysis strongly predicts a mitochondrial signal sequence in the extended N-terminal of human rTSgamma, which is absent in rTSbeta. We confirmed the existence of rTS in human mitochondria experimentally by demonstrating the presence of both rTSgamma and rTSbeta proteins in mitochondria isolated by subcellular fractionation. In addition, our comprehensive analysis of rTS orthologous sequences reveals an unusual phylogenetic distribution of this gene, which suggests the occurrence of one or more horizontal gene transfer events. CONCLUSION: The presence of two rTS isoforms in mitochondria suggests that the rTS signaling pathway may be active within mitochondria. Our report also presents an example of identifying novel protein isoforms and for improving gene annotation through comparative genomic analysis.

Amino Acid Sequence↗

Chromosome-Scale Genome Analysis Reveals Locus-Specific Disruption of the Citrinin-Associated Region in a Furu-Derived Monascus ruber Strain BC20.

Monascus species are widely used in traditional fermented foods for pigment and flavor formation, but citrinin contamination remains a major safety concern that limits broader food applications. Therefore, this study aimed to evaluate the citrinin risk of a furu-derived Monascus ruber strain, BC20, by integrating phenotypic screening across food-relevant matrices with genome-resolved analysis. After 14 days of cultivation across eight matrices, including fungal media as well as dairy-, cereal-, and bran-based substrates, citrinin was not detected by immunoaffinity cleanup combined with HPLC-FLD (LOD, 4 μg/kg; LOQ, 12 μg/kg). To investigate the genetic basis of this phenotype, we generated a chromosome-scale genome assembly for BC20 and conducted comparative analyses across a total of 19 Monascus genomes. ANI analysis and phylogenomic inference consistently placed BC20 within the ruber-pilosus clade. Comparative synteny analysis showed that the citrinin-associated locus in BC20 no longer retained an intact cluster configuration but instead exhibited a remnant-locus architecture, and similar patterns were also observed in several related genomes from the same clade. By contrast, the monacolin K (mk) locus remained syntenically conserved in BC20, supporting locus-specific structural disturbance rather than assembly-derived pseudo-absence. Additionally, its antifungal susceptibility was determined. Overall, BC20 represents a M. ruber candidate strain with undetectable citrinin, and this study provides a practical analytical framework for citrinin risk screening in food-related Monascus isolates.

biosynthetic gene cluster↗

Comparative genome analysis and pathway reconstruction.

Pathway reconstruction builds on genome and biochemical data with the aim of reconstructing higher level interactions between identified enzymes in a specific genome, in particular the different enzyme pathways (species or individual/patient). Metabolite flow in a pathway is analyzed by different tools, such as elementary mode analysis. This reveals key enzymes and pharmacological targets in the enzyme network. An overview of bioinformatic tools and algorithms for these tasks, application examples and recent results from these techniques are presented. Target selection, drug development and optimization can all be sped up using these approaches.

Animals↗