Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Full-genome analysis of resistance gene homologues in rice.

The availability of the rice genome sequence enabled the global characterization of nucleotide-binding site (NBS)-leucine-rich repeat (LRR) genes, the largest class of plant disease resistance genes. The rice genome carries approximately 500 NBS-LRR genes that are very similar to the non-Toll/interleukin-1 receptor homology region (TIR) class (class 2) genes of Arabidopsis but none that are homologous to the TIR class genes. Over 100 of these genes were predicted to be pseudogenes in the rice cultivar Nipponbare, but some of these are functional in other rice lines. Over 80 other NBS-encoding genes were identified that belonged to four different classes, only two of which are present in dicotyledonous plant sequences present in databases. Map positions of the identified genes show that these genes occur in clusters, many of which included members from distantly related groups. Members of phylogenetic subgroups of the class 2 NBS-LRR genes mapped to as many as ten different chromosomes. The patterns of duplication of the NBS-LRR genes indicate that they were duplicated by many independent genetic events that have occurred continuously through the expansion of the NBS-LRR superfamily and the evolution of the modern rice genome. Genetic events, such as inversions, that inhibit the ability of recently duplicated genes to recombine promote the divergence of their sequences by inhibiting concerted evolution.

Amino Acid Sequence↗

Genome analysis of human rotaviruses by oligonucleotide mapping of isolated RNA segments.

The genomes of adult diarrhoea rotaviruses isolated in different parts of China during winter outbreaks in 1983 and 1984 were compared by segmental oligonucleotide (ON) mapping. The RNA profiles of most of the isolates were indistinguishable but it was found that some corresponding RNA segments had identical or very closely related ON maps whereas others differed considerably. This finding can be taken to suggest that the strains compared may be genetically related by a natural reassortment event. The genomes of cocirculating group A rotaviruses isolated in Scotland during winter outbreaks in 1981/82 were also compared. The ON maps of corresponding RNA segment differed extensively irrespective of whether or not the segments comigrated on gels.

Adult↗

Genome analysis of marine photosynthetic microbes and their global role.

Four recently completed genome projects on marine Cyanobacteria have started the age of comparative genomics for marine microbes. Cyanobacteria are a group of photoautotrophic bacteria that have traditionally been under-represented in studies of complete genome sequences, as have microbes from the marine environment in general. The new genome information is of crucial importance to understanding their role in oceanic primary production, global carbon cycling and functioning of the biosphere. Marine microbes are a still almost untapped resource for the identification of novel beneficial metabolites and activities. The availability of an increasing number of genome sequences will eventually lead to a sustained development of marine biotechnology.

Bacteriophages↗

Whole-genome analysis: annotations and updates.

The most important advances in the field of genome annotation over the past two years involve the use of cDNA sequences, protein structures and gene expression data to predict genes. These types of information not only improve gene identification, but they also give insights into variation in gene structure and function.

Alternative Splicing↗

A genomic analysis of two-component signal transduction in Streptococcus pneumoniae.

A genomics-based approach was used to identify the entire gene complement of putative two-component signal transduction systems (TCSTSs) in Streptococcus pneumoniae. A total of 14 open reading frames (ORFs) were identified as putative response regulators, 13 of which were adjacent to genes encoding probable histidine kinases. Both the histidine kinase and response regulator proteins were categorized into subfamilies on the basis of phylogeny. Through a systematic programme of mutagenesis, the importance of each novel TCSTS was determined with respect to viability and pathogenicity. One TCSTS was identified that was essential for the growth of S. pneumoniaeThis locus was highly homologous to the yycFG gene pair encoding the essential response regulator/histidine kinase proteins identified in Bacillus subtilis and Staphylococcus aureus. Separate deletions of eight other loci led in each case to a dramatic attenuation of growth in a mouse respiratory tract infection model, suggesting that these signal transduction systems are important for the in vivo adaptation and pathogenesis of S. pneumoniae. The identification of conserved TCSTSs important for both pathogenicity and viability in a Gram-positive pathogen highlights the potential of two-component signal transduction as a multicomponent target for antibacterial drug discovery.

Animals↗

Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.

We present here the annotation of the complete genome of rice Oryza sativa L. ssp. japonica cultivar Nipponbare. All functional annotations for proteins and non-protein-coding RNA (npRNA) candidates were manually curated. Functions were identified or inferred in 19,969 (70%) of the proteins, and 131 possible npRNAs (including 58 antisense transcripts) were found. Almost 5000 annotated protein-coding genes were found to be disrupted in insertional mutant lines, which will accelerate future experimental validation of the annotations. The rice loci were determined by using cDNA sequences obtained from rice and other representative cereals. Our conservative estimate based on these loci and an extrapolation suggested that the gene number of rice is approximately 32,000, which is smaller than previous estimates. We conducted comparative analyses between rice and Arabidopsis thaliana and found that both genomes possessed several lineage-specific genes, which might account for the observed differences between these species, while they had similar sets of predicted functional domains among the protein sequences. A system to control translational efficiency seems to be conserved across large evolutionary distances. Moreover, the evolutionary process of protein-coding genes was examined. Our results suggest that natural selection may have played a role for duplicated genes in both species, so that duplication was suppressed or favored in a manner that depended on the function of a gene.

Arabidopsis↗

Single-read sequence tags of a limited number of genomic DNA fragments provide an inexpensive tool for comparative genome analysis.

Single-read sequences from both ends of 415 3-kb average size genomic DNA fragments of Candida albicans were compared with the complete sequence data of Saccharomyces cerevisiae. Comparison at the protein level, translated DNA against protein sequences, revealed 138 sequence tags with clear similarity to S. cerevisiae proteins or open reading frames. One case of synteny was found for the open reading frames of RAD16 and LYS2, which are adjacent to each other in S. cerevisiae and C. albicans.

Adenosine Triphosphatases↗

Complete genome analysis of RFLP 184 isolates of porcine reproductive and respiratory syndrome virus.

Two full-length genomes of recently emerged virulent isolates of porcine reproductive and respiratory syndrome virus (PRRSV) were sequenced and compared to other PRRSV strains. The results revealed that these two isolates (named MN184), of North American lineage, represented the shortest PRRSV genomes sequenced to date with a nucleotide length of 15019 bases. Genetic analysis demonstrated that the two isolates were not identical and shared approximately 87 and 59% nucleotide identity with prototype North American strain VR-2332 and European strain Lelystad, respectively. Three quite variable regions were identified, corresponding to putative nsp1beta, putative nsp2 and ORF5. Nsp2, the most variable region, shared only 66-70% amino acid similarity to other sequenced North American-like PRRSV nsp2 proteins. Further study revealed that the nsp2 protein of the MN184 isolates contained three discontinuous deletions when compared to strain VR-2332 nsp2 protein, with the sizes of 111, 1, and 19 amino acids corresponding to strain VR-2332 positions 324-434, 486 and 505-523, respectively. The results suggest that targeted manipulation of PRRSV through nsp2 modification by reverse genetics may yield promising vectors for vaccine development, as has been recently demonstrated [Han, J., Faaberg, K.S., Wang, Y., Liu, H., 2005. Non-structural protein 2 mutants of PRRSV strain VR-2332 infectious clone based on deletions seen in RFLP184 isolates are viable. In: PRRS International Symposium Proceedings, vol. 8, Saint Louis, MO].

Amino Acid Sequence↗

Genomic analysis of breed composition and population structure in Montana composite cattle.

The Montana composite was developed in Brazil from crosses between Bos indicus and Bos taurus and structured into four biological types: Zebu (N), adapted taurine (A), British taurine (B), and continental taurine (C). This study aimed to characterize the genetic diversity and population structure of the Montana composite using genomic data through principal component analysis (PCA), admixture analysis, and Wright's FST statistic. The PCA revealed a clear separation between Bos indicus and Bos taurus groups, with Montana animals distributed in an intermediate position. The first two principal components explained 69.48% and 3.45% of the total variation, respectively. Supervised admixture estimates indicated a predominance of taurine contribution, with type A accounting for 34.47%, 52.64%, and 51.71% at K&#x2009;=&#x2009;4, 9, and 11, respectively. Increasing the ancestry resolution refined the contribution of individual founder breeds without changing the overall predominance of taurine ancestry. Comparisons between breed proportions obtained from pedigree and genomic data revealed significant differences, for most biological types and ancestry models (P&#x2009;<&#x2009;0.001), indicating that realized breed composition deviates from theoretical expectations. Estimates of genetic differentiation confirmed greater divergence between Zebu and taurine groups, as well as reduced distances among populations sharing common ancestry. Specific relationships were identified between the composite and some of its founder breeds, particularly Belmont Red, Senepol, and Tuli. Overall, the results demonstrate that the Montana composite has a complex genomic structure, with genomic ancestry varying according to the resolution adopted and differing from pedigree-based expectations.

Animals↗

SPC: a SPectral Component approach leveraging Identity-by-Descent graphs to address recent population structure in genomic analysis.

Population structure is a well-known confounder in statistical genetics, particularly in genome-wide association studies (GWAS), where it can lead to inflated test statistics and spurious associations. Traditional methods, such as principal components (PCs), commonly used to adjust for population structure, are limited in capturing fine-scale, non-linear patterns that arise from recent demographic events - patterns that are crucial for understanding rare variant effects. To address this challenge, we propose a novel method called SPectral Components (SPCs), which leverages identity-by-descent (IBD) graphs to capture and transform local, non-linear fine-scale population structure into continuous representations that can be seamlessly integrated into genetic analysis pipelines. Using both simulated datasets and empirical data from the UK Biobank (N &#x2248; 420,000), we demonstrate that SPCs outperform PCs in adjusting for fine-scale population structure. In simulations, SPCs explained over 90% of the fine-scale population structure with fewer components, while PCs captured less than 5%. In the UK Biobank, SPCs reduced the inflation of p-values in the GWAS of an environmental-driven phenotype by 12% compared to PCs, while maintaining a similar performance to PCs in height, a highly heritable phenotype. Additionally, SPCs improved rare variant association analyses, reducing genomic inflation (e.g., from 7.6 to 1.2 in one analysis), and provided more accurate heritability estimates. Spatial autocorrelation analysis further confirmed the ability of SPCs to account for environmental effects, reducing Moran's I for both environmental and heritable phenotypes more effectively than PCs. Overall, our findings demonstrate that SPCs provide a robust, scalable adjustment for recent population structure, offering a powerful alternative or complement to PCs in large-scale biobank studies.

GWAS↗

Genomic analysis of Mycobacterium tuberculosis complex strains used for production of purified protein derivative.

The genomes of the tuberculin production strains Mycobacterium bovis AN5 and Mycobacterium tuberculosis DT were compared to genome-sequenced tubercle bacilli by using DNA microarrays. Neither the AN5 nor DT strain suffered extensive gene deletions during in vitro passage. This suggests that bovine tuberculin made from M. bovis AN5 is suitable to detect infection with presently prevalent M. bovis strains.

Gene Deletion↗

Genomic analysis of carbon source metabolism of Shewanella oneidensis MR-1: Predictions versus experiments.

Genomic sequences have been used to find the genetic foundation for carbon source metabolism in Shewanella oneidensis MR-1. Annotated S. oneidensis MR-1 gene products were examined for their sequence similarity to enzymes participating in pathways for utilization of carbon and energy as described in the BioCyc database (http://www.biocyc.org/) or in the primary literature. A picture emerges that relegates five- and six-carbon sugars to minor roles as carbon sources, whereas multiple pathways for utilization of up to three-carbon carbohydrates seem to be present. Capacity to utilize amino acids for carbon and energy is also present. A few contradictions emerged in which enzymes appear to be present by annotations but are not active in the cell according to physiological experiments. Annotations are based on close sequence similarity and will not reveal inactivity due to deleterious mutations or due to lack of coordination of regulation and transport. Genes for a few enzymes known by experiment to be active are not found in the genome. This may be due to extensive divergence after duplication or convergence of function in separate lines in evolution rendering activities undetectable by sequence similarity. To minimize false predictions from protein sequences, we have been conservative in predicting pathways. We did not predict any pathway when, although a partial pathway was seen it was composed largely of enzymes already accounted for in any other complete pathway. This is an example of how a biochemically oriented sequence analysis can generate questions and direct further experimental investigation.

Bacterial Proteins↗

CGHAnalyzer: a stand-alone software package for cancer genome analysis using array-based DNA copy number data.

SUMMARY: This synopsis provides an overview of array-based comparative genomic hybridization data display, abstraction and analysis using CGHAnalyzer, a software suite, designed specifically for this purpose. CGHAnalyzer can be used to simultaneously load copy number data from multiple platforms, query and describe large, heterogeneous datasets and export results. Additionally, CGHAnalyzer employs a host of algorithms for microarray analysis that include hierarchical clustering and class differentiation. AVAILABILITY: CGHAnalyzer, the accompanying manual, documentation and sample data are available for download at http://acgh.afcri.upenn.edu. This is a Java-based application built in the framework of the TIGR MeV that can run on Microsoft Windows, Macintosh OSX and a variety of Unix-based platforms. It requires the installation of the free Java Runtime Environment 1.4.1 (or more recent) (http://www.java.sun.com).

Algorithms↗

Full genomic analysis of hepatitis delta virus prevalent on Miyako Island, Japan.

The aims of this study were to determine the full-length genome sequences of hepatitis delta virus (HDV) in HDV RNA-positive subjects, and to elucidate the molecular specificity of the HDVs that are clustered on a distant island in Japan. This study included 3 subjects with chronic hepatitis who were positive for hepatitis B surface (HBs) antigen and HDV RNA, and who were admitted to the Okinawa Prefectural Miyako Hospital in 1998. The full-length genome sequence of HDV was determined by nested polymerase chain reaction (PCR) using four kinds of primer sets. The genomic length of HDV was 1,675, 1,679 and 1,681 base pairs, respectively. There was 90-92% nucleotide homology between each pair of isolates. In comparison with HDV isolates in geographically neighboring regions, the nucleotide homology of the 3 HDV isolates were 73-75% with the China isolate of genotype I; 77-78% with the Taiwan isolate of genotype I; 83-84% with a Japan isolate of genotype IIa; 85-87% with the Taiwan isolate of genotype IIa, and 87-88% with the Taiwan isolates of genotype IIb. Therefore, the Miyako Island isolates had high homology with the Taiwan isolate of genotype IIa and IIb. Phylogenetic analysis of the full-length genome sequences of HDV revealed that the two Miyako Island isolates were classified into a genotype IIb'. The other one was classified as genotype IIb. In conclusion, the HDV of Miyako Island isolates can be classified as a novel subgroup of genotype IIb, designated type IIb', and genotype IIb.

Aged↗

Lessons learned from the genome analysis of ralstonia solanacearum.

Ralstonia solanacearum is a devastating plant pathogen with a global distribution and an unusually wide host range. This bacterium can also be free-living as a saprophyte in water or in the soil in the absence of host plants. The availability of the complete genome sequence from strain GMI1000 provided the basis for an integrative analysis of the molecular traits determining the adaptation of the bacterium to various environmental niches and pathogenicity toward plants. This review summarizes current knowledge and speculates on some key bacterial functions, including metabolic versatility, resistance to metals, complex and extensive systems for motility and attachment to external surfaces, and multiple protein secretion systems. Genome sequence analysis provides clues about the evolution of essential virulence genes such as those encoding the Type III secretion system and related pathogenicity effectors. It also provided insights into possible mechanisms contributing to the rapid adaptation of the bacterium to its environment in general and to its interaction with plants in particular.

Amino Acid Sequence↗