Search PubMedSearch

SEARCH · Search PubMed

Results for “Terminal Repeat Sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The chromosome-level genome of Stylosanthes guianensis provides insights into genome evolution and environmental adaptation.

Stylosanthes guianensis is a leguminous forage crop of significant economic importance, primarily distributed in tropical and subtropical regions. It exhibits strong adaptability to various stresses, yet the genetic basis underlying this trait remains unclear. In this study, we constructed the first chromosome-scale reference genome of S. guianensis using a combination of Nanopore and Hi-C sequencing technologies. The assembled genome size is 1254 Mb, with 10 pseudochromosomes. Using Nanopore full-length transcriptome data, we generated high-quality transcript-level gene annotations, identifying 36 585 gene models and 110 601 transcripts. The repetitive sequences in S. guianensis account for 79.16% of the genome, with the extensive expansion of Gypsy elements in long terminal repeats contributing to its genome size enlargement. Comparative genomic and transcriptomic analyses revealed that flavonoid metabolism plays a pivotal role in stress adaptation, providing new insights into the genetic basis of stress tolerance. Additionally, we generated whole-genome methylation profiles under cold treatment and control conditions, offering valuable data for future epigenomic research. These findings provide essential molecular resources for understanding stress resilience in S. guianensis and advancing its molecular breeding.

Genome, Plant

Structure of Herpesvirus saimiri genomes: arrangement of heavy and light sequences in the M genome.

Herpesvirus saimiri contains two species of DNA molecules. (i) The M genome is composed of 70% light (L) DNA (36% cytosine plus guanine; density in CsCl, 1.695 g/ml), which consists of unique sequences, and 30% heavy (H) DNA (71% cytosine plus guanine; density, 1.729 g/ml). (ii) The H genome contains heavy sequences exclusively. H sequences in M and H genomes cross-hybridize completely and are cleaved identically by restriction endonuclease R-Sma I into four classes of fragments with molecular weights of about 360,000, 300,000, 130,000 and 40,000, respectively. H sequences are chains of identical repeat units in tandem arrangement. The molecular weight of each repeat unit is about 830,000. L sequences have no cleavage site for endo R-Sma I H sequences are terminally arranged at both ends of the M genome, as seen by electron microscopy after partial denaturation. The length of the individual heavy ends varies between 21 mum and less than 1 mum, whereas the light region is uniform in size (35.3+/-0.35 mum). As a rule, molecules with a long heavy end at one side have a short heavy end at the other side, thus giving rise to a limited size heterogeneity. Orientation of M DNA molecules by the denaturation map of the light region shows that the longer heavy end may be located at the left or at the right side of the M genome.

Base Sequence

Membrane penicillinase of Bacillus licheniformis 749/C:sequence and possible repeated tetrapeptide structure of the phospholipopeptide region.

The membrane penicillinase (EC 3.5.2.6; penicillin amido-beta-lactamhydrolase) of Bacillus licheniforis 749/C, which appears to be an intermediate in the formation of the exoenzyme, is a phospholipoprotein that carries an NH2-terminal chain of 24 amino acids (only serine, glycine, aspartic acid, asparagine, glutamic acid, and glutamine) and a phosphatidylserine that is not present in the exoenzyme.

Amino Acid Sequence

A novel allele of Sh1 facilitates the development of waxy-sweet corn from waxy corn.

Waxy corn and sweet corn represent 2 major classes of fresh-eating corn, each with distinct sensory attributes and nutritional compositions. Developing a new variety that combines both waxy and sweet traits would address rising consumer demand and expand new market potential. From a fast neutron-mutagenized population of the waxy corn inbred line HB522, we isolated a novel mutant, designated as wx-sweet, whose kernels simultaneously exhibit waxy and sweet characteristics at the milk-filling stage. Through bulked segregant analysis combined with fine mapping, we mapped the causal locus to SHRUNKEN1 (Sh1) on chromosome 9, which was confirmed by an allelism test with a characterized Mu-insertion allele of Sh1. A 7,227-bp Copia-type long terminal repeat retrotransposon insertion was identified in exon 2 of Sh1 in the wx-sweet mutant by long-read sequencing. Consistently, the novel sh1 allele significantly reduced sucrose synthase activity. Genetic and physiological analyses demonstrate that sh1 and wx1 act synergistically to fine-tune carbohydrate metabolism in the endosperm. Integrated transcriptomic and metabolomic profiling uncover extensive transcriptional reprogramming and redirected metabolic flux, leading to substantial accumulation of sucrose and a range of oligosaccharides. These metabolic shifts underlie the unique simultaneous dual waxy-sweet texture in fresh-eating wx-sweet kernels. In summary, our work not only provides valuable genetic resources for breeding next-generation fresh-eating corn but also, for the first time, elucidates the molecular mechanism by which the sh1 and wx1 mutations cooperatively shape the waxy-sweet endosperm phenotype.

Zea mays

Complete covalent structure of statherin, a tyrosine-rich acidic peptide which inhibits calcium phosphate precipitation from human parotid saliva.

The complete amino acid sequence of human salivary statherin, a peptide which strongly inhibits precipitation from supersaturated calcium phosphate solutions, and therefore stabilizes supersaturated saliva, has been determined. The NH2-terminal half of this Mr=5380 (43 amino acids) polypeptide was determined by automated Edman degradations (liquid phase) on native statherin. The peptide was digested separately with trypsin, chymotrypsin, and Staphylococcus aureus protease, and the resulting peptides were purified by gel filtration. Manual Edman degradations on purified peptide fragments yielded peptides that completed the amino acid sequence through the penultimate COOH-terminal residue. These analyses, together with carboxypeptidase digestion of native statherin and of peptide fragments of statherin, established the complete sequence of the molecule. The 2 serine residues (positions 2 and 3) in statherin were identified as phosphoserine. The amino acid sequence of human salivary statherin is striking in a number of ways. The NH2-terminal one-third is highly polar and includes three polar dipeptides: H2PO3-Ser-Ser-H2PO3-Arg-Arg-, and Glu-Glu-. The COOH-terminal two-thirds of the molecule is hydrophobic, containing several repeating dipeptides: four of -Gn-Pro-, three of -Tyr-Gln-, two of -Gly-Tyr-, two of-Gln-Tyr-, and two of the tetrapeptide sequence -Pro-Tyr-Gln-Pro-. Unusual cleavage sites in the statherin sequence obtained with chymotrypsin and S. aureus protease were also noted.

Amino Acid Sequence

Mapping of regions on cloned Saccharomyces cerevisiae 2-mum DNA coding for polypeptides synthesized in Escherichia coli minicells.

Saccharomyces cerevisiae 2-mum DNA and some of its restriction fragments were integrated in vector pCR1 ,pBR313 or pBR322 and their expression in Escherichia coli P678-54 minicells was analyzed. 2mum DNA inserted at the EcoRI site of pCR1 or pBR313 and at the PstI site of pBR322 promoted the synthesis of polypeptides of 48,000, 37,000, 35,000 and 19,000 daltons. The DNA regions coding for these polypeptides were mapped on the 2-mum DNA molecule by insertion of single EcoRI or HindIII restriction fragments and comparison of the polypeptides produced. For the synthesis of the 37,000 dalton polypeptide, intact sites RIB and H3 were required. The disappearance of the 37,000 dalton polypeptide on interruption of one of these sites by insertion of the vector, was correlated with the appearance of a polypeptide of 22,000 or 23,500 daltons respectively. The DNA sequence coding for the 37,000 daltons polypeptide, therefore, has to be located in the S-loop region close to or overlapping with the site RIB and H3. Assuming that the 22,000 and the 23,500 dalton polypeptides are truncated forms of the 37,000 dalton polypeptide, the last polypeptide can be exactly mapped. The polypeptide of 48,000 daltons was mapped to that half of the L-loop segment containing the sites H1 and H2. If, however, HindIII fragment H1-H2 was expressed, the 48,000 dalton polypeptide was lost and concomitantly a 43,000 dalton polypeptide appeared. We assume that this polypeptide results from early termination of the polypeptide of 48,000 daltons. The 35,000 and 9,000 dalton polypeptides were mapped to the S-loop region. The integrated inverted repeat sequence of yeast 2-mum Dna did not induce any detectable insert-specific polypeptide synthesis.

Base Sequence

Sequence of 1019 nucleotides encompassing one of the inverted repeats from the yeast 2 micrometer plasmid.

A sequence of 1019 nucleotides encompassing one of the 600 base inverted repeats and non-repeated flanking regions has been determined in the type A yeast 2 micrometers plasmid cloned in pMB9. Methods are described for applying the Maxam-Gilbert sequencing procedure to DNA fragments labelled at the 3'-end using a T4-polymerase exchange/repair reaction and for sequencing 5'-end labelled fragments using dideoxy-nucleotides as chain terminators in the presence of E. coli DNA polymerase (nach Klenow). A notable feature of the sequence is its unusual content of symmetry elements. In one region of 140 nucleotides, 137 are involved in a complex arrangement of direct and inverted repeats linked by palindromic sequences.

Base Sequence

Moloney murine sarcoma virions synthesize full-genome-length double-stranded DNA in vitro.

Moloney murine sarcoma virus (MSV) virions incubated under optimal conditions were shown to support extensive synthesis of double-stranded DNA. The major product, a 5950-base-pair (6-kilobase-pair DNA) double-stranded DNA, was characterized by cleavage with restriction endonucleases and shown to contain a 600-nucleotide-long direct repeat at both ends of the MSV genome. Linear DNA molecules made in vivo shortly after infection were compared to the linear double-stranded DNA synthesized in vitro. The restriction maps of both viral DNA products were indistinguishable. The 600-base-pair repeat results in a progeny DNA molecule that is longer than the parental MSV genomic RNA. The generation of this repeat must involve a mechanism that allows the viral reverse transcriptase (RNA-dependent DNA nucleotidyltransferase) to copy 5'- and 3'-terminal genomic (+) strand sequences twice.

Chromosome Mapping

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant

Nucleotide sequence of the Hind-C fragment of simian virus 40 DNA. Comparison of the 5'-untranslated region of wild-type virus and of some deletion Mutants.

We report here the nucleotide sequence of the wild-type simian virus 40 (strain 776) restriction fragment Hind-C-P1 DNA and of the homologous region of various mutant DNAs which lack part of this fragment. During this work, we detected between EcoRII fragments N and G an additional, 17-base-pair EcoRII fragment, fragment P, which had previously been overlooked. Also, an additional dTpdG dinucleotide at residues L 339--340 was observed by sequence analysis of the DNA minus (E) strand; the presence of this dinucleotide was masked on sequencing patterns of the plus strand due to the persistence (during gel electrophoresis) of some secondary structures in the strand's 5'-terminal region. These nucleotide additions raise the total length of SV40 DNA to 5243 base pairs. The longest tandemly repeated segment in SV40 DNA now extends over 72 base pairs. SV40 deletion mutants dl 893 and dl 894 and SV40 strains Rh 911 and 1801 all lack an identical 72-base-pair-long DNA segment in the Hind-C region. This deletion corresponds precisely to one of the two aforementioned large tandemly repeated sequences. Mutant dl 895 lacks 66 base pairs, 63 of which are part of the former repetition. All these mutants, except dl 895, very probably were generated by an intramolecular, homologous recombination event. The 40-base-pair deletion in mutant dl 1811 includes the major capping site of SV40 late RNA. dl 1812 lacks only three base pairs, which are part of the overlapping HhaI and HpaII restriction sites at position 0.725--0.726.

Base Sequence

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta

Cloned single repeating units of 5S DNA direct accurate transcription of 5S RNA when injected into Xenopus oocytes.

Single and multiple repeating units of three types of Xenopus 5S DNA recombined with the plasmid pMB9 serve as templates for the accurate synthesis of 5S RNA after their injection into Xenopus laevis oocyte nuclei. All 15 cloned single repeating units of X. laevis oocyte 5S DNA that were tested supported 5S RNA synthesis. Three cloned fragments of X. borealis oocyte 5S DNA and one cloned single repeating unit of X. borealis somatic 5S DNA were templates for 5S RNA synthesis. We conclude that the majority of repeating units of 5S DNA in these multigene families contain the information for accurate initiation and termination of 5S RNA synthesis. The ability of this system to detect sequence changes that affect transcription is demonstrated.

Animals

The repeating nucleotide sequence in the repetitive mitochondrial DNA from a "low-density" petite mutant of yeast.

The repeating nucleotide sequence of 68 base pairs in the mtDNA from an ethidium-induced cytoplasmic petite mutant of yeast has been determined. For sequence analysis specifically primed and terminated RNA copies, obtained by in vitro transcription of the separated strands, were use. The sequence consists of 66 consecutive AT base pairs flanked by two GC pairs and comprises nearly all of the mutant mitochondrial genome. The sequence, moreover, also represents the first part of wild-type mtDNA sequence so far.

Base Sequence

On inverted repeat sequences in chromosomal DNA.

It is suggested that chromosomal DNA should contain a class of palindromic reverse repeats, comparable in number to that of genes themselves, which are formed as follows:(1) a transcription-termination signal that follows the gene plus (on the complementary strand and located as near to the "anti-gene" as possible); (2) a second termination signal which actively prevents the accidental transcription of the anti-gene. Thus, the adjacent termination-anti-termination region of one strand would complement the anti-termination-termination region of the other.

Chromosome Inversion

Isolation and characterization of a heptaglycosylceramide from bovine erythrocyte membranes.

A heptaglycosylceramide was isolated from bovine erythrocyte membranes. The structure was characterized to be Gal(alpha 1-3)Gal(beta 1-4)GlcNAc(beta1-3)Gal(beta 1-4)Glc-NAc(beta 1-4)al(beta 1-4)GlcCer. A hexaglycosylceramide that has the same sequence except for the terminal alpha-galactosyl unit has also been isolated. We have previously found that gangliosides isolated from bovine erythrocyte membranes contain a keratan sulfate type repeating unit --[3Gal(beta 1-4)-GlcNAc beta]--n. This study shows that the keratan sulfate type repeating unit is also present in the neutral glycosphingolipids of bovine erythrocyte membranes.

Animals

Studies on the primary structure of bovine high-molecular-weight kininogen. Amino acid sequence of a fragment ("histidine-rich peptide") released by plasma kallikrein.

An unknown peptide fragment, which was released from bovine high-molecular-weight kininogen by bovine plasma kallikrein [EC 3.4.21.8], was isolated and its chemical structure was established. The fragment consisted of 41 amino acids with serine and arginine at the NH2- and COOH-termini, respectively. The molecular weight was calculated to be 4,584. It was very basic and contained eleven residues each of histidine and glycine and seven residues of lysine. Thus, the total number of these three amino acids accounted for about 70 percent of the total residues constituting the fragment. The amino acid sequence of the fragment, designated tentatively as "His-rich peptide," was studied by Edman degradation and standard enzymatic and chemical techniques. These data made it possible to deduce the following sequence: H-Ser-His-Gly-Leu-Gly-His-Gly-His-Gln-Lys-Gln-His-Gly-Leu-Gly-His-Gly-His-Lys-His-Gly-His-Gly-His-Gly-Lys-His-Lys-Asn-Lys-Gly-Lys-Asn-Asn-Gly-Lys-His-Tyr-Asp-Trp-Arg-OH. The fragment had an extremely interesting feature in that repeating sequences occur along the peptide chain. The repeats were of the type His-Gly-X or Gly-His-X and this sequence appeared six or seven times up to 26 residues from the N-terminal end. Moreover, three tetrapeptide sequences of Gly-His-Gly-His and two heptapeptide sequence consisting of His-Gly-Leu-Gly-His-Gly-His were found in the N-terminal portion. It should be noted that plasma kallikrein liberates such a histidine-rich peptide from the kininogen in addition to a physiologically active peptide, bradykinin. The location of the "His-rich peptide" fragment in the percursor protein is also discussed.

Amino Acid Sequence

Pea histones H2A and H2B. Variable and conserved regions in the sequences.

Pea histone II group, a mixture of H2A and H2B obtained by chromatography on an ion-exchange resin, was further fractionated by carboxymethylcellulose chromatography and purified by Bio-Gel P-60 chromatography. Their chromatographic behaviors and gel electrophoretic mobilities of single bands differed significantly from those of calf H2A and H2B. Their amino acid compositions were similar to those of the calf histones as a whole, but differed in detail in certain respects. The partial sequence of pea H2B was deduced from the amino acid compositions of BrCN cleavage fragments and tryptic peptides in comparison with the known sequence of calf H2B. It is different in the amino-terminal basic region from the calf H2B, with a blocked amino terminal and a larger number of residues. In contrast, the middle and carboxy-terminal hydrophobic regions are relatively similar, with at least 19-21 different residues and microheterogeneity at two positions of the pea sequence. The sequence of H2A may vary in much the same way as that of H2B, as suggested by the similar extent of differences in their amino acid compositions. It is thus assumed that the amino-terminal regions, at least, of H2A and H2B histones are variable in evolution provided that they remain basic enough to bind DNA, whereas the middle and carboxy-terminal hydrophobic regions of H2A, H2B, H3, and H4 should be conserved to ensure precise histone core formation inside the repeated units of chromatin.

Amino Acid Sequence

The polysaccharides from heterocyst and spore envelopes of a blue-green alga. Structure of the basic repeating unit.

The polysaccharides from the envelopes of heterocysts and spores of Anabaena cylindrica consist of repeating units containing 1 mannosyl and 3 glucosyl residues, all linked by beta(1 yields 3) glycosidic bonds, with glycosidic bonds, with glucose, xylose, galactose, and mannose present in side branches. Degradation of the polysaccharides with specific glycosidases has permitted identification of the linkages to almost all of the branches. When the polysaccharides, from which all but two types of side branches had been cleaved, were digested with a beta(1 yields 3) endoglucanase, glucose, a tri-, and a pentasaccharide were produced. The oligosaccharide products were identified as (see article of journal). The backbones of the polysaccharides were sequenced from the reducing terminus by a modified Smith degradation. Analysis with NaB3H4 at each stage of the degradation showed that the backbones terminate in the sequence Man-Glc-Glc-Glc and are therefore presumed to have the structure (Man-Glc-Glc-Glc)n, and that they contain an average of from 128 to 150 sugar residues. From the information obtained, the repeating sequences of the original polysaccharides from the two types of differentiated cells of A. cylindrica could be largely deduced and appeared to be identical.

Carbohydrates