Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genome composition”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Antigenic characteristics and genome composition of a naturally occurring recombinant influenza virus isolated from a pig in Japan.

We performed antigenic analysis of the haemagglutinin and neuraminidase subunits of a recombinant virus (A/swine/Kanagawa/2/78) isolated from a pig in Japan in 1978, using a series of monoclonal antibodies to H1 (Hsw1) haemagglutinin and N2 neuraminidases of H2N2 and H3N2 viruses. Results obtained in haemagglutination inhibition tests with five monoclonal antibodies to the haemagglutinin of A/NJ/8/76 (H1N1) revealed that the haemagglutinin of three H1N1 and the recombinant viruses were indistinguishable from that of A/NJ/8/76. The neuraminidase of A/swine/Kanagawa/2/78 was found to be antigenically similar to A/Kumamoto/22/76 (H3N2, A/Victoria/3/75-like strain). The oligonucleotide maps of the entire RNAs of H1N1, H1N2 and H3N2 viruses showed that A/swine/Kanagawa/2/78 (H1N2) virus was more similar to swine (H1N1) virus than to A/Kumamoto/22/76 (H3N2) virus. Radioactive cDNA was prepared by reverse transcription of the recombinant virus RNA using a dodecadeoxyribonucleotide primer and used in DNA-RNA hybridization experiments. The results obtained in molecular hybridization based on blotting procedures showed that all cDNA segments except gene 6 hybridized efficiently with RNAs of swine (H1N1) influenza virus. The sixth cDNA segment was homologous to the corresponding RNA segment of H3N2 virus. The genetic relatedness of A/swine/Kanagawa/2/78 (H1N2) with either A/swine/Kanagawa/4/78 (H1N1) or A/Kumamoto/22/76 (H3N2) was clearly established by hybridization between the cDNA segment probes and viral RNA. It was concluded that the neuraminidase gene of A/swine/Kanagawa/2/78 (H1N2) was derived from a human H3N2 virus, while the seven other genes were from a swine H1N1 virus.

Animals↗

The composite genome of the legume symbiont Sinorhizobium meliloti.

The scarcity of usable nitrogen frequently limits plant growth. A tight metabolic association with rhizobial bacteria allows legumes to obtain nitrogen compounds by bacterial reduction of dinitrogen (N2) to ammonium (NH4+). We present here the annotated DNA sequence of the alpha-proteobacterium Sinorhizobium meliloti, the symbiont of alfalfa. The tripartite 6.7-megabase (Mb) genome comprises a 3.65-Mb chromosome, and 1.35-Mb pSymA and 1.68-Mb pSymB megaplasmids. Genome sequence analysis indicates that all three elements contribute, in varying degrees, to symbiosis and reveals how this genome may have emerged during evolution. The genome sequence will be useful in understanding the dynamics of interkingdom associations and of life in soil environments.

Bacterial Adhesion↗

A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.

BACKGROUND: Correlations between genome composition (in terms of GC content) and usage of particular codons and amino acids have been widely reported, but poorly explained. We show here that a simple model of processes acting at the nucleotide level explains codon usage across a large sample of species (311 bacteria, 28 archaea and 257 eukaryotes). The model quantitatively predicts responses (slope and intercept of the regression line on genome GC content) of individual codons and amino acids to genome composition. RESULTS: Codons respond to genome composition on the basis of their GC content relative to their synonyms (explaining 71-87% of the variance in response among the different codons, depending on measure). Amino-acid responses are determined by the mean GC content of their codons (explaining 71-79% of the variance). Similar trends hold for genes within a genome. Position-dependent selection for error minimization explains why individual bases respond differently to directional mutation pressure. CONCLUSIONS: Our model suggests that GC content drives codon usage (rather than the converse). It unifies a large body of empirical evidence concerning relationships between GC content and amino-acid or codon usage in disparate systems. The relationship between GC content and codon and amino-acid usage is ahistorical; it is replicated independently in the three domains of living organisms, reinforcing the idea that genes and genomes at mutation/selection equilibrium reproduce a unique relationship between nucleic acid and protein composition. Thus, the model may be useful in predicting amino-acid or nucleotide sequences in poorly characterized taxa.

Amino Acids↗

Hierarchical structure analysis describing abnormal base composition of genomes.

Abnormal base compositional patterns of genomic DNA sequences are studied in the framework of a hierarchical structure (HS) model originally proposed for the study of fully developed turbulence [She and Lévêque, Phys. Rev. Lett. 72, 336 (1994)]. The HS similarity law is verified over scales between 10(3)bp and 10(5)bp, and the HS parameter beta is proposed to describe the degree of heterogeneity in the base composition patterns. More than one hundred bacteria, archaea, virus, yeast, and human genome sequences have been analyzed and the results show that the HS analysis efficiently captures abnormal base composition patterns, and the parameter beta is a characteristic measure of the genome. Detailed examination of the values of beta reveals an intriguing link to the evolutionary events of genetic material transfer. Finally, a sequence complexity (S) measure is proposed to characterize gradual increase of organizational complexity of the genome during the evolution. The present study raises several interesting issues in the evolutionary history of genomes.

Animals↗

Molecular cytogenetic analyses of hexaploid lines spontaneously appearing in octoploid Triticale.

Genome characterization of 14 hexaploid lines that spontaneously appeared in octoploid Triticales was carried out by sequential genomic in situ hybridization and fluorescence in situ hybridization, high molecular weight glutenin subunits and SSR marker analyses. All of the lines showed a chromosome constitution of complete A and B genomes, and a composite genome consisting of the chromosomes of D and R genomes. The composite genome of the 11 lines consisted of chromosomes 1R, 2D, 3R, 4R, 5R, 6R and 7R, that of the two lines were 1D, 2D, 3R, 4R, 5R, 6R and 7R, and that of one line was 1R, 2D, 3R, 4R, 5R, 6D and 7R. The incompatibility of the D and R genomes in common wheat genetic background, preferential retention of chromosome 2D and importance of these lines for the development of hexaploid Triticale are discussed in this report.

Chromosomes, Plant↗

Compositional correlation studies among the three different codon positions in 12 bacterial genomes.

Compositional distributions in the three codon positions of the coding sequences of 12 fully sequenced prokaryotic genomes, which are publicly available, were investigated. A universal compositional correlation was observed in most of the genomes under investigation irrespective of their overall genomic GC contents. In all the genomes, the GC contents at the first codon positions are always greater than the overall GC contents of the genomes whereas the reverse is true in the case of second codon positions. GC contents at the third codon positions are higher than the overall genomic GC contents in high GC containing genomes, and the opposite situation was found in case of low GC genomes except for Helicobacter pylori. In high-GC rich genomes, the GC contents at the first + second codon positions are less than the GC contents at the third codon positions, and they are low in low-GC genomes except for Helicobacter pylori. The distributions of four bases at the three different positions were also investigated for all 12 organisms. It was observed that in high-GC genomes G is the most dominant base and in low-GC genomes A is the most dominant base in the first codon positions. But purine bases, i.e., (A + G), predominantly occur in the first codon position. In the second codon position, A is the most dominant base in most of the organisms and G is the least dominant base in all the organisms. There is no unique regular pattern of individual bases at the third codon positions; however, there are significant differences in the occurrences of (G + C) contents in the third codon positions among the different organisms. Calculations of dinucleotide frequencies in 12 different organisms indicate that in GC-rich genomes GG, GC, CC, and CG dinucleotides are the most dominant whereas the reverse is true in case of low-GC genomes. Biological implications of these results are discussed in this paper.

Bacteria↗

A molecular genome scan analysis to identify chromosomal regions influencing economic traits in the pig. I. Growth and body composition.

Genome scans can be employed to identify chromosomal regions and eventually genes (quantitative trait loci or QTL) that control quantitative traits of economic importance. A three-generation resource family was developed by using two Berkshire grand sires and nine Yorkshire grand dams to detect QTL for growth and body composition traits in pigs. A total of 525 F2 progeny were produced from 65 matings. All F2 animals were phenotyped for birth weight, 16-day weight, growth rate, carcass weight, carcass length, back fat thickness, and loin eye area. Animals were genotyped for 125 microsatellite markers covering the genome. Least squares regression interval mapping was used for QTL detection. All carcass traits were adjusted for live weight at slaughter. A total of 16 significant QTL, as determined by a permutation test, were detected at the 5% chromosome-wise level for growth traits on Chromosomes (Chrs) 1, 2, 3, 4, 6, 7, 8, 9, 11, 13, 14, and X, of which two were significant at the 5% genome-wise level and two at the 1% genome-wise level (on Chrs 1, 2, and 4). For composition traits, 20 QTL were significant at the 5% chromosome-wise level (on Chrs 1, 4, 5, 6, 7, 12, 13, 14, 18), of which one was significant at the 5% genome-wise level and three were significant at the 1% genome-wise level (on Chrs 1, 5, and 7). For several QTL the favorable allele originated from the breed with the lower trait mean.

Animals↗

A mathematical method for determining genome divergence and species delineation using AFLP.

The delineation of bacterial species is presently achieved using direct DNA-DNA relatedness studies of whole genomes. It would be helpful to obtain the same genomically based delineation by indirect methods, provided that descriptions of individual genome composition of bacterial genomes are obtained and included in species descriptions. The amplified fragment length polymorphism (AFLP) technique could provide the necessary data if the nucleotides involved in restriction and amplification are fundamental to the description of genomic divergences. Firstly, in order to verify that AFLP analysis permits a realistic exploration of bacterial genome composition, the strong correspondence between predicted and experimental AFLP data was demonstrated using Agrobacterium strain C58 as a model system. Secondly, a method is proposed for determining current genome mispairing and evolutionary genome divergences between pairs of bacteria, based on arbitrary sampling of genomes by using AFLP. The measure of current genome mispairing was validated by comparison with DNA-DNA relatedness data, which itself correlates with base mispairing. The evolutionary genome divergence is the estimated rate of nucleotide substitution that has occurred since the strains diverged from a common ancestor. Current genome mispairing and evolutionary genome divergence were used to compare members of Agrobacterium, used as a model of closely related genomic species. A strong and highly significant correlation was found between calculated genome mispairing and DNA-DNA relatedness values within genomic species. The canonical 70% DNA-DNA hybridization value used to delineate genomic species was found to correspond to a range of current genome mispairing of 13-13.6%. These limits correspond to 0.097 and 0.104 nucleotide substitutions per site, respectively. In addition, experimental data showed that the large Ti and cryptic plasmids of Agrobacterium had little effect on the estimation of genome divergence. Evolutionary genome divergence was used for phylogenetic inferences. Data showed that members of the same genomic species clustered consistently, as supported by bootstrap resampling. On the basis of these results, it is proposed that the genomic delineation of bacterial species could be based, in future, on phylogenetic groups supported by bootstraps and genome descriptions of individual strains, obtained by AFLP analysis, recorded in accessible databases; this approach might eventually replace DNA-DNA hybridization studies.

Biological Evolution↗

Investigating the relationship between genome structure, composition, and ecology in prokaryotes.

Our thesis is that the DNA composition and structure of genomes are selected in part by mutation bias (GC pressure) and in part by ecology. To illustrate this point, we compare and contrast the oligonucleotide composition and the mosaic structure in 36 complete genomes and in 27 long genomic sequences from archaea and eubacteria. We report the following findings (1) High-GC-content genomes show a large underrepresentation of short distances between G(n) and C(n) homopolymers with respect to distances between A(n) and T(n) homopolymers; we discuss selection versus mutation bias hypotheses. (2) The oligonucleotide compositions of the genomes of Neisseria (meningitidis and gonorrhoea), Helicobacter pylori and Rhodobacter capsulatus are more biased than the other sequenced genomes. (3) The genomes of free-living species or nonchronic pathogens show more mosaic-like structure than genomes of chronic pathogens or intracellular symbionts. (4) Genome mosaicity of intracellular parasites has a maximum corresponding to the average gene length; in the genomes of free-living and nonchronic pathogens the maximum occurs at larger length scales. This suggests that free-living species can incorporate large pieces of DNA from the environment, whereas for intracellular parasites there are recombination events between homologous genes. We discuss the consequences in terms of evolution of genome size. (5) Intracellular symbionts and obligate pathogens show small, but not zero, amount of chromosome mosaicity, suggesting that recombination events occur in these species.

AT Rich Sequence↗

A comprehensive analysis of mammalian mitochondrial genome base composition and improved phylogenetic methods.

Phylogenetic analysis of mammalian species using mitochondrial protein genes has proved to be problematic in many previous studies. The high mutation rate of mitochondrial DNA and unusual base composition of several species has prompted us to conduct a detailed study of the composition of 69 mammalian mitochondrial genomes. Most major changes in base composition between lineages can be attributed to shifts between the proportions of C and T on the L-strand. These changes are significant at all codon positions and are shown to affect amino acid composition. Correlated changes in the base composition of the RNA loops and stems are also observed. Following up from previous studies, we investigate changes in the base composition of all 12 H-strand proteins and find that variability in proportions of C and T is correlated with location on the genome. Variation in base composition across genes and species is known to adversely affect the performance of phylogenetic inference methods. We have, therefore, developed a customized three-state general time-reversible DNA substitution model, implemented in the PHASE phylogenetic inference package, which lumps C and T into a composite pyrimidine state. We compare the phylogenetic tree obtained using the new three-state model with that obtained using a standard four-state model. Results using the three-state model are more congruent with recent studies using large sets of nuclear genes and help resolve some of the apparent conflicts between studies using nuclear and mitochondrial proteins.

Animals↗

Intimate evolution of proteins. Proteome atomic content correlates with genome base composition.

Discerning the significant relations that exist within and among genome sequences is a major step toward the modeling of biopolymer evolution. Here we report the systematic analysis of the atomic composition of proteins encoded by organisms representative of each kingdoms. Protein atomic contents are shown to vary largely among species, the larger variations being observed for the main architectural component of proteins, the carbon atom. These variations apply to the bulk proteins as well as to subsets of ortholog proteins. A pronounced correlation between proteome carbon content and genome base composition is further evidenced, with high G+C genome content being related to low protein carbon content. The generation of random proteomes and the examination of the canonical genetic code provide arguments for the hypothesis that natural selection might have driven genome base composition.

Animals↗

Correlation between codon usage, regional genomic nucleotide composition, and amino acid composition in the cytochrome P-450 gene superfamily.

The codon usage bias of 110 mammalian cytochrome P-450 genes has been determined and analyzed in relation to a variety of genetic, biochemical, and physiological parameters. In those P-450 genes exhibiting biased usage the preferred codons generally do not differ among the four species examined (rat, rabbit, man, and mouse) or from the predominantly used codons identified for all sequenced genes in a recent data base analysis (Wada et al. (1992) Nucleic Acids Res. 20 (Suppl.), 2111-2118). Codon usage bias does not correlate with evolutionary relationships, evolutionary age, or with the extent of evolutionary conservation of orthologous proteins; there is no obvious correlation with the level of expression of a given P-450, with its inducibility, nor with its physiologic role; and neither the preferred codons nor the degree of bias differ for P-450s expressed in different tissues. Codon usage bias does correlate with the C+G content at the codon third position, and thus preferred codons usually end in C or G; for those P-450s for which gene sequences are available this bias also correlates with the C + G content of the intronic and flanking regions of these genes. Moreover, a lesser increase in the C + G content at the codon first and second positions is also evident in genes located in regions of high C + G content; this leads to predictable differences in the amino acid compositions of P-450 enzymes that correlate with genomic nucleotide composition and the degree of bias in codon usage.

Amino Acids↗

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial↗

CVTree: a phylogenetic tree reconstruction tool based on whole genomes.

Composition Vector Tree (CVTree) implements a systematic method of inferring evolutionary relatedness of microbial organisms from the oligopeptide content of their complete proteomes (http://cvtree.cbi.pku.edu.cn). Since the first bacterial genomes were sequenced in 1995 there have been several attempts to infer prokaryote phylogeny from complete genomes. Most of them depend on sequence alignment directly or indirectly and, in some cases, need fine-tuning and adjustment. The composition vector method circumvents the ambiguity of choosing the genes for phylogenetic reconstruction and avoids the necessity of aligning sequences of essentially different length and gene content. This new method does not contain 'free' parameter and 'fine-tuning'. A bootstrap test for a phylogenetic tree of 139 organisms has shown the stability of the branchings, which support the small subunit ribosomal RNA (SSU rRNA) tree of life in its overall structure and in many details. It may provide a quick reference in prokaryote phylogenetics whenever the proteome of an organism is available, a situation that will become commonplace in the near future.

Algorithms↗

Amino acid composition of genomes, lifestyles of organisms, and evolutionary trends: a global picture with correspondence analysis.

Can we infer the lifestyle of an organism from the characteristic properties of its genome? More precisely, what are the relations between easily quantifiable properties from genomic sequences, such as amino-acid compositions, and more subtle characteristics concerning for example lifestyles or evolutionary trends? Here, we seek a global picture for such properties, based on a large number (56) of complete genomes, including significant numbers of representatives from the three domains of life. We consider the amino acid compositions of the predicted proteomes, and we use correspondence analysis, as a multivariate method to extract the relevant information from the large-scale data. From these analyses we derive a series of conclusions, concerning lifestyles, as well as physico-chemical and evolutionary trends: (1) correspondence analysis of the amino acid compositions permits discrimination between the three known lifestyles (mesophily/thermophily/hyperthermophily). (2) For various organisms, amino-acid composition properties are essentially driven by GC content, and to a significantly lesser extent by growth temperatures associated with lifestyles. Roughly speaking, the respective contributions of these two components are 57 and 20%. It is notable that these proportions are essentially unchanged with respect to a previous analysis (Nature 393 (1998) 537), which involved only 15 genomes, available at the time. (3) In terms of amino acid compositional biases, two specific 'signatures' for thermophily (in a broad sense, including hyperthermophily) can be detected. First, thermophilic species display a relative abundance in glutamic acid (Glu), concomitantly with the depletion in glutamine. Second, in thermophilic species, the relative abundance in Glu (negative charge) is significantly correlated (Pearson correlation coefficient r=0.83 with P<0.0001), with the increase in the lumped 'pool' lysine+arginine (positive charges). This correlation (absent in mesophiles) could be interpreted on a physico-chemical basis, relevant to the thermostability of proteins. (4) Statistically significant differences are observed between the average lengths of the genes in the surveyed species, which follow their distribution between the three domains of life. Also a significant difference is observed between the average lengths of thermophilic (283.0+/-5.8) versus mesophilic (340+/-9.4) genes. It is thus possible that the 'general' shortening of the primary sequences in thermophilic proteins plays a role in thermostability. (5) Considering various combinations of conservation properties (genes conserved exclusively in eukaryotes, in archaea, in bacteria, in combinations of two domains, etc.) correspondence analysis reveals a trend towards thermophilic-hyperthermophilic profiles for the most conserved subset of genes (ancient genes). (6) When limited to the subset of species-specific genes, correspondence analysis leads to a different picture for the clustering of genomes following amino-acid compositions: for example, the 'core' specific part of a genome can bear lifestyle signatures different from those of the complete genome.Various results are discussed both on methodological and biological grounds. The evolutionary perspectives opened by our analyses are noted.

Amino Acids↗

The compositional evolution of vertebrate genomes.

The compositional evolution of vertebrate genomes is characterized: (i) by one predominant conservative mode, in which nucleotide changes occur, but the base composition of DNA sequences in general, and of coding sequences in particular, does not change; and (ii) by three different shifting or transitional modes, in which nucleotide changes are accompanied by changes in the base composition of sequences. Investigations on these evolutionary modes have shed new light on a central problem in molecular evolution, namely the role played by natural selection in modulating the mutational input. This review will present first the intragenomic shifts, the 'major shifts' and the 'minor shift', and then the 'whole-genome', or 'horizontal', shift. In each case, the shifts were preceded and followed by a conservative mode of evolution. This review expands on a previous one [Bernardi, Gene 241 (2000) 3-17], and summarizes the evidence that the changes of the compositional patterns of the genome and their maintenance are controlled by Darwinian natural selection.

Animals↗

Is there replication-associated mutational pressure in the Saccharomyces cerevisiae genome?

Compositional bias of yeast chromosomes was analysed using detrended DNA walks. Unlike eubacterial chromosomes, the yeast chromosomes did not show the specific asymmetry correlated with origin and terminus of replication. It is probably a result of a relative excess of autonomously replicating sequences (ARS) and of random choice of these sequences in each replication cycle. Nevertheless, the last ARS from both ends of chromosomes are responsible for unidirectional replication of subtelomeric sequences with pre-established leading/lagging roles of DNA strands. In these sequences a specific asymmetry is observed, resembling the asymmetry introduced by replication-associated mutational pressure into eubacterial chromosomes.

Animals↗