Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

PCR-based gene synthesis as an efficient approach for expression of the A+T-rich malaria genome.

The A+T-rich genome of the human malaria parasite Plasmodium falciparum encodes genes of biological importance that cannot be expressed efficiently in heterologous eukaryotic systems, owing to an extremely biased codon usage and the presence of numerous cryptic polyadenylation sites. In this work we have optimized an assembly polymerase chain reaction (PCR) method for the fast and extremely accurate synthesis of a 2.1 kb Plasmodium falciparum gene (pfsub-1) encoding a subtilisin-like protease. A total of 104 oligonucleotides, designed with the aid of dedicated computer software, were assembled in a single-step PCR. The assembly was then further amplified by PCR to produce a synthetic gene which has been cloned and successfully expressed in both Pichia pastoris and recombinant baculovirus-infected High Five(TM) cells. We believe this strategy to be of special interest as it is simple, accessible and has no limitation with respect to the size of the gene to be synthesized. Used as a systematic approach for the malarial genome or any other A + T-rich organism, the method allows the rapid synthesis of a nucleotide sequence optimized for expression in the system of choice and production of sufficiently large amounts of biological material for complete molecular and structural characterization.

Amino Acid Sequence↗

Analysis of the hexon gene sequence of bovine adenovirus type 4 provides further support for a new adenovirus genus (Atadenovirus).

The putative hexon gene of bovine adenovirus type 4 (BAV-4), encoding 910 amino acid residues, has been identified and sequenced. A characteristic codon usage biased towards the use of AT-rich triplets was observed. Comparative analysis with other hexon sequences detected a high level of amino acid identity in the regions corresponding to the pedestals of the hexon. Substitutions, insertions and deletions were identified mainly in the variable regions forming the loops which are exposed on the outer surface of the virion. In these variable regions, BAV-4 shared similarity only with egg drop syndrome (EDS) virus and ovine adenovirus isolate 287 (OAV287). The close relationship of these viruses was also demonstrated by phylogenetic analysis of the hexon gene. In addition to the two groups of the Mastadenovirus and Aviadenovirus genera, a third cluster appeared comprising BAV-4, OAV287 and EDS virus.

Amino Acid Sequence↗

The nucleotide sequence of a streptomycin streptomycin phosphotransferase (streptomycin kinase) [corrected] gene from a streptomycin producer.

The nucleotide sequence of the DNA fragment containing the streptomycin phosphotransferase (streptomycin kinase) [corrected] gene from the streptomycin-producer Streptomyces griseus strain HUT 6037 was determined. Analysis of the sequence revealed an open reading frame which could encode 325 amino acid residues. A biased codon usage pattern, reflecting the high G + C composition (approximately 74%) of Streptomyces DNA, was observed in the gene.

Amino Acid Sequence↗

Cloning of a second non-haem bromoperoxidase gene from Streptomyces aureofaciens ATCC 10762: sequence analysis, expression in Streptomyces lividans and enzyme purification.

The gene for BPO-A1, one of two non-haem bromoperoxidases in the tetracycline and 7-chlorotetracycline producer Streptomyces aureofaciens ATCC 10762, was cloned in the positive selection vector pIJ699 and expressed in Streptomyces lividans TK64. The cloned bromoperoxidase was over-produced up to 2800-fold by the S. lividans TK64 transformant. By taking advantage of the over-production of BPO-A1 and the heat stability of the enzyme, a new and simple purification procedure was developed. Subcloning into the vector pIJ487 and screening of recombinants by a newly developed histochemical assay located the bpoA1 gene on a 2.1 kb BamHI-HindIII fragment. The nucleotide sequence of the 2.1 kb fragment was determined; the bpoA1 gene was identified within the sequence on the basis of the biased codon usage of Streptomyces genes and the presence of a nucleotide sequence encoding the N-terminal amino acid sequence obtained from the purified BPO-A1. Comparison of the deduced primary structure of BPO-A1 with those deduced for the non-haem chloroperoxidase CPO-P from Pseudomonas pyrrocinia and the bromoperoxidase BPO-A2 from S. aureofaciens ATCC 10762 gave amino acid sequence identities of 49% and 40%, respectively.

Amino Acid Sequence↗

Hierarchy of sequence-dependent features associated with prokaryotic translation.

Protein expression in the cell is affected by various sequence-dependent features. Several such sequence-dependent features have been individually studied,yet they have not been compared quantitatively in terms of their relative influence on protein expression,and a hierarchy of these elements has not been determined. Here we present a quantitative analysis examining sequence-dependent features involved in prokaryotic translation,namely,the base-pairing potential between the mRNA Shine-Dalgarno sequence and the ribosomal RNA,codon bias,and the identity of the stop codon. We analyzed these features both at intra- and intergenomic levels using the Escherichia coli and Haemophilus influenzae genomes. Within each genome,we examined the relationship between each feature and protein expression levels determined by 2D-gel analyses. At the intergenomic level,comparative genomic principles were applied to study the relative preservation of the different sequence-dependent properties between orthologs. From these analyses,we determined that biased codon usage is the property that is most highly associated with protein expression and that is most conserved. The identity of the stop codon and the base-pairing potential of the mRNA Shine-Dalgarno sequence and the rRNA seem to have less of an effect on protein expression.

3' Untranslated Regions↗

PmSUC3: characterization of a SUT2/SUC3-type sucrose transporter from Plantago major.

Higher plants possess medium-sized gene families that encode plasma membrane-localized sucrose transporters. For several plant species, it has been shown that at least one of these genes (e.g., AtSUC3 in Arabidopsis and LeSUT2 in tomato) differs from all other family members in several features, such as the length of the open reading frame, the number of introns, and the codon usage bias. For these reasons, and because two of these proteins did not rescue a yeast mutant defective in sucrose utilization, it had been speculated that this subgroup of transporters might have sensor functions. Here, we describe the detailed functional characterization and cellular localization of PmSUC3, the orthologous transporter from the Plantago major transporter family. The PmSUC3 protein is localized in the sieve elements of the Plantago phloem and mediates the energy-dependent transport of sucrose and maltose. In contrast to the situation in solanaceous plants, PmSUC3 is not colocalized with PmSUC2, the source-specific, phloem-loading sucrose transporter of Plantago. Moreover, PmSUC3 also was identified in sieve elements of sink leaves and in several nonphloem cells and tissues. Arguments for and against a potential sensor function for this type of sucrose transporter are presented, and the role of this type of transporter in the regulation of sucrose fluxes is discussed.

Amino Acid Sequence↗

Determinants of translational initiation efficiency in the atp operon of Escherichia coli.

Transcription and translation of the atp genes encoding the subunits b, delta, alpha, gamma and epsilon of the Escherichia coli H+-ATPase were studied. The nature and quantities of the respective transcripts initiated from different promoters were compared with overall expression rates thus yielding accurate information about relative translational efficiency and its coupling to mRNA levels. Part of the highly efficient subunit c gene translational initiation region (TIR) was used as a tool in manipulating the TIRs of the other genes. Rate control of atp cistron translation occurs at the initiation level and is determined locally by each gene's TIR. In this way, individual subunit synthesis rates are set to match the requirements for H+-ATPase assembly. There is no (or very restricted) translational coupling between the cistrons. Translational initiation rates of the normally weakly expressed atp genes could be increased by up to a factor of 27 by manipulating the sequences upstream of the start codons, despite biased codon usages. In the presence of an improved upstream sequence, the N-terminal sequence of the subunit gamma gene exerted a limiting effect. This could be relieved by altering the sequence of the first seven codons. The levels of subunit gamma mRNA were more sensitive to changes in translational efficiency than the concentrations of the other atp mRNAs. The relationships between initiation efficiency and primary and secondary structure in the natural and manipulated atp TIRs are discussed in detail.

Base Sequence↗

Cloning and characterization of the isopenicillin N synthase gene of Streptomyces griseus NRRL 3851 and studies of expression and complementation of the cephamycin pathway in Streptomyces clavuligerus.

A gene, pcbC, encoding the isopenicillin N synthase of Streptomyces griseus NRRL 3851, has been cloned in a 6.4-kb Bg/II DNA fragment and located in an internal 1.55-kb PvuII segment by hybridization with the Penicillium chrysogenum pcbC gene. Hybridization studies revealed the presence of homologous sequences in the DNAs of several Streptomyces strains and Nocardia lactamdurans. The S. griseus pcbC gene was not expressed in Streptomyces lividans but was expressed in Streptomyces clavuligerus and complemented a mutation, nce2, that impaired isopenicillin N synthase and cephamycin biosynthesis. The pcbC gene contained an open reading frame of 990 nucleotides that encodes a protein of 329 amino acids with a deduced Mr of 37,371. The isopenicillin N synthase formed after expression of the pcbC gene in the S. clavuligerus nce2 mutant strain was found to have an Mr of 38,000 by gel filtration. A protein of about 38 kDa was observed in sodium dodecyl sulfate-polyacrylamide gel electrophoresis gels of extracts of a transformant of the nce2 mutant strain; this protein was absent from the untransformed mutant strain. The G+C content of the pcbC gene was 63.6%, and the strongly biased codon usage was typical of that of Streptomyces strains. A transcription initiation site was found 44 nucleotides upstream of the ATG translation initiation triplet. A transcript of 1.1 kb was observed in the donor S. griseus strain and also in the S. clavuligerus nce2 mutant strain transformed with the pcbC gene, suggesting that it is transcribed as a monocistronic mRNA.

Actinomycetales↗

Bifidobacterium lactis DSM 10140: identification of the atp (atpBEFHAGDC) operon and analysis of its genetic structure, characteristics, and phylogeny.

The atp operon is highly conserved among eubacteria, and it has been considered a molecular marker as an alternative to the 16S rRNA gene. PCR primers were designed from the consensus sequences of the atpD gene to amplify partial atpD sequences from 12 Bifidobacterium species and nine Lactobacillus species. All PCR products were sequenced and aligned with other atpD sequences retrieved from public databases. Genes encoding the subunits of the F(1)F(0)-ATPase of Bifidobacterium lactis DSM 10140 (atpBEFHAGDC) were cloned and sequenced. The deduced amino acid sequences of these subunits showed significant homology with the sequences of other organisms. We identified specific sequence signatures for the genus Bifidobacterium and for the closely related taxa Bifidobacterium lactis and Bifidobacterium animalis and Lactobacillus gasseri and Lactobacillus johnsonii, which could provide an alternative to current methods for identification of lactic acid bacterial species. Northern blot analysis showed that there was a transcript at approximately 7.3 kb, which corresponded to the size of the atp operon, and a transcript at 4.5 kb, which corresponded to the atpC, atpD, atpG, and atpA genes. The transcription initiation sites of these two mRNAs were mapped by primer extension, and the results revealed no consensus promoter sequences. Phylogenetic analysis of the atpD genes demonstrated that the Lactobacillus atpD gene clustered with the genera Listeria, Lactococcus, Streptococcus, and Enterococcus and that the higher G+C content and highly biased codon usage with respect to the genome average support the hypothesis that there was probably horizontal gene transfer. The acid inducibility of the atp operon of B. lactis DSM 10140 was verified by slot blot hybridization by using RNA isolated from acid-treated cultures of B. lactis DSM 10140. The rapid increase in the level of atp operon transcripts upon exposure to low pH suggested that the ATPase complex of B. lactis DSM 10140 was regulated at the level of transcription and not at the enzyme assembly step.

Bacterial Proteins↗

Cloning and sequencing of the immunoglobulin A1 protease gene (iga) of Haemophilus influenzae serotype b.

Secretion of immunoglobulin A1 (IgA1) proteases is a characteristic of Haemophilus influenzae and several other bacterial pathogens causing infectious diseases, including meningitis. Indirect evidence suggests that the proteases are important virulence factors. In this study, we cloned the iga gene encoding immunoglobulin A1 (IgA1) protease from H. influenzae serotype b into Escherichia coli, in which the recombinant H. influenzae iga gene was expressed and the resulting protease was secreted. Sequencing a part of a 7.5-kilobase DNA fragment containing the iga gene revealed a large open reading frame with a strongly biased codon usage and having the potential of encoding a protein of 1,541 amino acids and a molecular mass of 169 kilodaltons. Putative promoter and terminator elements flanking the open reading frame were identified. Comparison of the deduced amino acid sequence of this H. influenzae IgA1 protease with that of a similar protease from Neisseria gonorrhoeae revealed several domains with a high degree of homology. Analogous to mechanisms known from the N. gonorrhoeae IgA protease secretion, we propose a scheme of posttranslational modifications of the H. influenzae IgA1 protease precursor, leading to a secreted protease with a molecular mass of 108 kilodaltons, which is close to the 100 kilodaltons reported for the mature IgA1 protease.

Amino Acid Sequence↗

Characterization of the Streptococcus pneumoniae immunoglobulin A1 protease gene (iga) and its translation product.

Bacterial immunoglobulin A1 (IgA1) proteases constitute a very heterogenous group of extracellular endopeptidases which specifically cleave human IgA1 in the hinge region. Here we report that the IgA1 protease gene, iga, of Streptococcus pneumoniae is homologous to that of Streptococcus sanguis. By using the S. sanguis iga gene as hybridization probe, the corresponding gene from a clinical isolate of S. pneumoniae was isolated in an Escherichia coli lambda phage library. A lysate of E. coli infected with hybridization-positive recombinant phages possessed IgA1-cleaving activity. The complete sequence of the S. pneumoniae iga gene was determined. An open reading frame with a strongly biased codon usage and having the potential of encoding a protein of 1,927 amino acids with a molecular mass of 215,023 Da was preceded by a potential -10 promoter sequence and a putative Shine-Dalgarno sequence. A putative signal peptide was found in the N-terminal end of the protein. The amino acid sequence similarity to the S. sanguis IgA1 protease indicated that the pneumococcal IgA1 protease is a Zn-metalloproteinase. The primary structures of the two streptococcal IgA1 proteases were quite different in the N-terminal parts, and both proteins contained repeat structures in this region. Using a novel assay for IgA1 protease activity upon sodium dodecyl sulfate-polyacrylamide gel electrophoresis, we demonstrated that the secreted IgA1 protease was present in several different molecular forms ranging in size from approximately 135 to 220 kDa. In addition, interstrain differences in the sizes of the pneumococcal IgA1 proteases were detected. Southern blot analyses suggested that the S. pneumoniae iga gene is highly heterogenous within the species.

Amino Acid Sequence↗

Cloning, sequencing, and characterization of the principal acid phosphatase, the phoC+ product, from Zymomonas mobilis.

The Zymomonas mobilis gene encoding acid phosphatase, phoC, has been cloned and sequenced. The gene spans 792 base pairs and encodes an Mr 28,988 polypeptide. This protein was identified as the principal acid phosphatase activity in Z. mobilis by using zymograms and was more active with magnesium ions than with zinc ions. Its promoter region was similar to the -35 "pho box" region of the Escherichia coli pho genes as well as the regulatory sequences for Saccharomyces cerevisiae acid phosphatase (PHO5). A comparison of the gene structure of phoC with that of highly expressed Z. mobilis genes revealed that promoters for all genes were similar in degree of conservation of spacing and identity with the proposed Z. mobilis consensus sequence in the -10 region. The phoC gene contained a 5' transcribed terminus which was AT rich, a weak ribosome-binding site, and less biased codon usage than the highly expressed Z. mobilis genes.

Acid Phosphatase↗

The cobalamin (coenzyme B12) biosynthetic genes of Escherichia coli.

The enteric bacterium Escherichia coli synthesizes cobalamin (coenzyme B12) only when provided with the complex intermediate cobinamide. Three cobalamin biosynthetic genes have been cloned from Escherichia coli K-12, and their nucleotide sequences have been determined. The three genes form an operon (cob) under the control of several promoters and are induced by cobinamide, a precursor of cobalamin. The cob operon of E. coli comprises the cobU gene, encoding the bifunctional cobinamide kinase-guanylyltransferase; the cobS gene, encoding cobalamin synthetase; and the cobT gene, encoding dimethylbenzimidazole phosphoribosyltransferase. The physiological roles of these sequences were verified by the isolation of Tn10 insertion mutations in the cobS and cobT genes. All genes were named after their Salmonella typhimurium homologs and are located at the corresponding positions on the E. coli genetic map. Although the nucleotide sequences of the Salmonella cob genes and the E. coli cob genes are homologous, they are too divergent to have been derived from an operon present in their most recent common ancestor. On the basis of comparisons of G+C content, codon usage bias, dinucleotide frequencies, and patterns of synonymous and nonsynonymous substitutions, we conclude that the cob operon was introduced into the Salmonella genome from an exogenous source. The cob operon of E. coli may be related to cobalamin synthetic genes now found among non-Salmonella enteric bacteria.

Bacterial Proteins↗

A computational approach for identifying pathogenicity islands in prokaryotic genomes.

BACKGROUND: Pathogenicity islands (PAIs), distinct genomic segments of pathogens encoding virulence factors, represent a subgroup of genomic islands (GIs) that have been acquired by horizontal gene transfer event. Up to now, computational approaches for identifying PAIs have been focused on the detection of genomic regions which only differ from the rest of the genome in their base composition and codon usage. These approaches often lead to the identification of genomic islands, rather than PAIs. RESULTS: We present a computational method for detecting potential PAIs in complete prokaryotic genomes by combining sequence similarities and abnormalities in genomic composition. We first collected 207 GenBank accessions containing either part or all of the reported PAI loci. In sequenced genomes, strips of PAI-homologs were defined based on the proximity of the homologs of genes in the same PAI accession. An algorithm reminiscent of sequence-assembly procedure was then devised to merge overlapping or adjacent genomic strips into a large genomic region. Among the defined genomic regions, PAI-like regions were identified by the presence of homolog(s) of virulence genes. Also, GIs were postulated by calculating G+C content anomalies and codon usage bias. Of 148 prokaryotic genomes examined, 23 pathogenic and 6 non-pathogenic bacteria contained 77 candidate PAIs that partly or entirely overlap GIs. CONCLUSION: Supporting the validity of our method, included in the list of candidate PAIs were thirty four PAIs previously identified from genome sequencing papers. Furthermore, in some instances, our method was able to detect entire PAIs for those only partial sequences are available. Our method was proven to be an efficient method for demarcating the potential PAIs in our study. Also, the function(s) and origin(s) of a candidate PAI can be inferred by investigating the PAI queries comprising it. Identification and analysis of potential PAIs in prokaryotic genomes will broaden our knowledge on the structure and properties of PAIs and the evolution of bacterial pathogenesis.

Bacteria↗

CGAT: a comparative genome analysis tool for visualizing alignments in the analysis of complex evolutionary changes between closely related genomes.

BACKGROUND: The recent accumulation of closely related genomic sequences provides a valuable resource for the elucidation of the evolutionary histories of various organisms. However, although numerous alignment calculation and visualization tools have been developed to date, the analysis of complex genomic changes, such as large insertions, deletions, inversions, translocations and duplications, still presents certain difficulties. RESULTS: We have developed a comparative genome analysis tool, named CGAT, which allows detailed comparisons of closely related bacteria-sized genomes mainly through visualizing middle-to-large-scale changes to infer underlying mechanisms. CGAT displays precomputed pairwise genome alignments on both dotplot and alignment viewers with scrolling and zooming functions, and allows users to move along the pre-identified orthologous alignments. Users can place several types of information on this alignment, such as the presence of tandem repeats or interspersed repetitive sequences and changes in G+C contents or codon usage bias, thereby facilitating the interpretation of the observed genomic changes. In addition to displaying precomputed alignments, the viewer can dynamically calculate the alignments between specified regions; this feature is especially useful for examining the alignment boundaries, as these boundaries are often obscure and can vary between programs. Besides the alignment browser functionalities, CGAT also contains an alignment data construction module, which contains various procedures that are commonly used for pre- and post-processing for large-scale alignment calculation, such as the split-and-merge protocol for calculating long alignments, chaining adjacent alignments, and ortholog identification. Indeed, CGAT provides a general framework for the calculation of genome-scale alignments using various existing programs as alignment engines, which allows users to compare the outputs of different alignment programs. Earlier versions of this program have been used successfully in our research to infer the evolutionary history of apparently complex genome changes between closely related eubacteria and archaea. CONCLUSION: CGAT is a practical tool for analyzing complex genomic changes between closely related genomes using existing alignment programs and other sequence analysis tools combined with extensive manual inspection.

Algorithms↗

Paleo-demography of the Drosophila melanogaster subgroup: application of the maximum likelihood method.

The species divergence times and demographic histories of Drosophila melanogaster and its three sibling species, D. mauritiana, D. simulans, and D. yakuba, were investigated using a maximum likelihood (ML) method. Thirty-nine orthologous loci for these four species were retrieved from DDBJ/EMBL/GenBank database. Both autosomal and X-linked loci were used in this study. A significant degree of rate heterogeneity across loci was observed for each pair of species. Most loci have the GC content greater than 50% at the third codon position. The codon usage bias in Drosophila loci is considered to result in the high GC content and the heterogenous rates across loci. The chi-square, G, and Fisher's exact tests indicated that data sets with 11, 23, and 9 pairs of DNA sequences for the comparison of D. melanogaster with D. mauritiana, D. simulans, and D. yakuba, respectively, retain homogeneous rates across loci. We applied the ML method to these data sets to estimate the DNA sequence divergences before and after speciation of each species pair along with their standard deviations. Using 1.6 x 10(-8) as the rate of nucleotide substitutions per silent site per year, our results indicate that the D. melanogaster lineage split from D. yakuba approximately 5.1 +/- 0.8 million years ago (mya), D. mauritiana 2.7 +/- 0.4 mya, and D. simulans 2.3 +/- 0.3 mya. It implies that D. melanogaster became distinct from D. mauritiana and D. simulans at approximately the same time and from D. yakuba no earlier than 10 mya. The effective ancestral population size of D. melanogaster appears to be stable over evolutionary time. Assuming 10 generations per year for Drosophila, the effective population size in the ancestral lineage immediately prior to the time of species divergence is approximately 3 x 10(6), which is close to that estimated for the extant D. melanogaster population. The D. melanogaster did not encounter any obvious bottleneck during the past 10 million years.

Animals↗

Gene expression intensity shapes evolutionary rates of the proteins encoded by the vertebrate genome.

Natural selection leaves its footprints on protein-coding sequences by modulating their silent and replacement evolutionary rates. In highly expressed genes in invertebrates, these footprints are seen in the higher codon usage bias and lower synonymous divergence. In mammals, the highly expressed genes have a shorter gene length in the genome and the breadth of expression is known to constrain the rate of protein evolution. Here we have examined how the rates of evolution of proteins encoded by the vertebrate genomes are modulated by the amount (intensity) of gene expression. To understand how natural selection operates on proteins that appear to have arisen in earlier and later phases of animal evolution, we have contrasted patterns of mouse proteins that have homologs in invertebrate and protist genomes (Precambrian genes) with those that do not have such detectable homologs (vertebrate-specific genes). We find that the intensity of gene expression relates inversely to the rate of protein sequence evolution on a genomic scale. The most highly expressed genes actually show the lowest total number of substitutions per polypeptide, consistent with cumulative effects of purifying selection on individual amino acid replacements. Precambrian genes exhibit a more pronounced difference in protein evolutionary rates (up to three times) between the genes with high and low expression levels as compared to the vertebrate-specific genes, which appears to be due to the narrower breadth of expression of the vertebrate-specific genes. These results provide insights into the differential relationship and effect of the increasing complexity of animal body form on evolutionary rates of proteins.

Animals↗

Mitogenome assembly and phylogenetic relationships of Phalaris arundinacea.

INTRODUCTION: As a perennial herb of Poaceae, Phalaris arundinacea plays key roles in grazing, production, and soil and water conservation because of its well-developed rhizomes and seed dispersal. We assembled and annotated the first mitogenome of P. arundinacea to support evolutionary and taxonomic research. METHODS: We assembled and annotated the first complete mitochondrial genome of P. arundinacea by integrating Illumina short reads with Nanopore long reads via a hybrid assembly strategy. The genome architecture was comprehensively characterized, encompassing codon usage bias, repetitive sequence organization, and inter-organellar genetic exchange with the chloroplast genome. RESULTS AND DISCUSSION: Assembly of the P. arundinacea mitogenome revealed two circular structures with a combined length of 526,717 bp. The genome comprised a set of 37 protein-coding genes (PCGs), 27 tRNAs, and 8 rRNAs, with the rRNA genes exhibiting full assembly (100% coverage). The mitochondrial genome contained 154 forward and 164 palindromic repeats, along with 25 tandem repeats and 124 simple sequence repeats (SSRs). Notably, 102 SSRs were distributed on contig1, predominantly in tetrameric form. Furthermore, 376 RNA editing sites were predicted. A total of 104 fragments were integrated into the mitochondrial genome from the chloroplast, amounting to 55,866 bp of transferred sequence. Finally, phylogenetic analysis of 28 plant mitogenomes placed P. arundinacea closest to species within the genus Poa (P. chaixii and P. pratensis). Comparative analysis of non-synonymous-to-synonymous substitution rate (Ka/Ks) ratios across divergent species revealed that the mitochondrial genome of P. arundinacea underwent stabilizing evolutionary dynamics, characterized by predominant purifying selection with several lineage-specific variations in selective pressure. Our findings support the close phylogenetic relationship between P. arundinacea and species of the genus Poa and provide a reference mitochondrial genome resource for future comparative studies within Phalaris that incorporate broader taxon sampling. These results support deeper phylogenetic investigations of P. arundinacea and facilitate future work on its germplasm characterization and applied use.

Phalaris arundinacea↗