Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Herpesvirus saimiri DNA in tumor cells--deleted sequences and sequence rearrangements.

Herpesvirus saimiri DNA in continuous lymphoblastoid cell lines obtained from viral induced tumors in marmosets has been analyzed by gel electrophoresis of restricted DNA. Southern transfer to nitrocellulose filters, and hybridization to 32P-labeled viral DNA or DNA fragments. The viral DNA fragments EcoRI-G, -H, -D, and -I, KpnI-A, and BamHI-D and -E were not detected in Southern transfers of DNA from the nonproducing 1670 cell line. For each restriction endonuclease, a new fragment appeared, consistent with a 13.0-megadalton deletion of viral DNA sequences. This deletion encompassed 35 to 48 +/- 0.6 megadaltons from the left end of the unique DNA region. A sequence arrangement map is presented for the major population of H. saimiri DNA sequences in the 1670 cell line. Although H. saimiri DNA in the nonproducing 70N2 cell line can be distinguished from viral DNA in the 1670 cell line by several criteria, the same sequences were found to be deleted in the major population of viral DNA molecules. Unlike 1670 and 70N2 cells, restricted DNA from the virus-producing cell lines 77/5 and 1926 contained all of the DNA fragments present in the parental virion DNA. DNA from 1670, 70N2, and 77/5 cells contained additional viral DNA fragments that did not comigrate with any virion DNA fragments. Most of these unexplained fragments were confined to or highly enriched in partially purified circular or linear DNA fractions. DNA from tumor cells taken directly from a tumor-bearing animal contained viral DNA indistinguishable from the parental virion DNA by the assay conditions used. These results indicate that viral DNA sequence rearrangements can occur upon cultivation of tumor cells in vitro and that excision of DNA sequences from the viral genome may play a role in establishing the nonproducing state of some tumor cell lines.

Animals↗

Efficient expression of the Saccharomyces cerevisiae PGK gene depends on an upstream activation sequence but does not require TATA sequences.

The Saccharomyces cerevisiae PGK (phosphoglycerate kinase) gene encodes one of the most abundant mRNA and protein species in the cell. To identify the promoter sequences required for the efficient expression of PGK, we undertook a detailed internal deletion analysis of the 5' noncoding region of the gene. Our analysis revealed that PGK has an upstream activation sequence (UASPGK) located between 402 and 479 nucleotides upstream from the initiating ATG sequence which is required for full transcriptional activity. Deletion of this sequence caused a marked reduction in the levels of PGK transcription. We showed that PGK has no requirement for TATA sequences; deletion of one or both potential TATA sequences had no effect on either the levels of PGK expression or the accuracy of transcription initiation. We also showed that the UASPGK functions as efficiently when in the inverted orientation and that it can enhance transcription when placed upstream of a TRP1-IFN fusion gene comprising the promoter of TRP1 fused to the coding region of human interferon alpha-2.

Base Sequence↗

Nucleotide sequence of the DNA encoding the 5'-terminal sequences of simian virus 40 late mRNA.

We have used a combination of techniques of DNA and RNA sequence analysis to determine the nucleotide sequence of the portion of simian virus 40 DNA preceding and encoding the 5' end of mRNA for the structural protein VP2 of simian virus 40. Comparison of the sequence with those found in polyadenylated RNA in the cytoplasm of infected cells RNA shows that the transcript of sequences preceding the structural gene is more abundant than the transcript containing the codons for the protein. Between the abundant transcript of sequences preceding the coding region and the less abundant transcript of the coding region there is a short sequence whose transcript is not detected.

Base Sequence↗

The amino acid sequence of rabbit skeletal alpha-tropomyosin. The NH2-terminal half and complete sequence.

The amino acid sequence of the large cyanogen bromide fragment (residues 11 to 127) derived from the NH2-terminal half of alpha-tropomyosin has been determined. This was achieved by automatic sequence analysis of the whole fragment as well as manual sequencing of fragments derived from tryptic digestion of the maleylated fragment and thermolytic, Myxobacter 495 alpha-lytic and Staphylococcus aureus protease digestion of the unmodified fragment. Methionine-containing overlap peptides have been isolated from tryptic digests of the maleylated protein as well as from S. aureus protease digests of the unmodified protein. Coupled with previously published information on the small cyanogen bromide fragments and methionine sequences of tropomyosin, these analyses have permitted the completion of the primary structure of the protein. The complete sequence differs by only 1 residue (Gln-24 instead of Glu-24) from that previously reported. Analysis of the sequence by several authors has permitted rational explanations for the stabilization of its coiled-coil structure, for the existence of its two chains in a nonstaggered arrangement, for a head-to-tail overlap of molecular ends of 8 to 9 residues, for the existence of 14 actin-binding sites on each tropomyosin molecule, and a suggestion for the site of binding of troponin-T.

Amino Acid Sequence↗

Pathways of transcript splicing in yeast mitochondria. Mutations in intervening sequences of the split gene COB reveal a requirement for intervening sequence-encoded products.

We have studied the transcript processing of the split gene COB (or BOX) in yeast mtDNA, in both wild type and cob- mutants. Using various DNA fragments specific for coding or intervening sequences of this gene, we have determined the composition of splicing intermediates by DNA/RNA hybridization. The pattern of splicing intermediates detected in wild type reveals differing rates of the five splicings resulting in an apparent pathway of processing rather than an absolute order among the five cut and splice events. Effects of mutations in four of the five sequences have been studied. All of them interfere with transcript processing. Some block the excision of the sequence mutated only, but allow other splicing events to occur essentially as in the wild type. They suggest that in these mutants any order of splicings is possible, but that some are preferred. In contrast, other mutations located in four different sequences block several splicings simultaneously and thus suggest the existence of an obligatory order of events. In order to reconcile these findings we discuss the following hypotheses. (i) Some intervening sequences in COB specify products which are involved in transcript splicing; (ii) the biosynthesis of trace amounts of these products occurs on splicing intermediates. Their formation requires a certain order of splicing events to occur on a small number of COB transcripts. (iii) If expressed and functional, the intervening sequence-encoded products, together with other components, act on the bulk of COB transcripts, resulting in the steady state pattern of splicing intermediates observed in wild type.

Base Sequence↗

Messenger ribonucleic acid of the lipoprotein of the Escherichia coli outer membrane. I. Nucleotide sequence at the 3' terminus and sequences of oligonucleotides derived from complete digests of the mRNA.

The sequence of 92 nucleotides at the 3' end of the mRNA which codes for the lipoprotein of the outer membrane of Escherichia coli has been determined to be GCUAACCAGCGUCUGGACAACAUGGCUACUAAAUACCGCAAGUAAUAGUACCUGUGAAGUGAAAAAUGGCGCACAUUGUGCGCCAUUUUUUUOH. This sequence includes the 50 nucleotides comprising the 3' untranslated region of the mRNA and contains codons for 14 amino acids at the COOH-terminal of the lipoprotein. In addition, the nucleotide sequences of all oligonucleotides derived from complete ribonuclease T1 and ribonuclease A digestions of the lipoprotein mRNA were established. These oligonucleotides were assigned to portions of the known amino acid sequence as well as the 5' untranslated and 3' untranslated regions of the mRNA molecule. With the use of the genetic code, these oligonucleotide sequences served to establish 94% of the mRNA sequence. The lipoprotein mRNA can be deduced to be 322 nucleotides in length. All three translation termination codons (UAA, UAG, and UGA) were found in phase with the coding region of the mRNA. The region at the 3' end of the mRNA showed unusual resistance to partial degradation, and the partial fragments from this region had anomalous mobilities in two-dimensional gels, even under denaturing conditions. This indicates that there is a very stable hairpin stem-and-loop structure at the 3' end. This hairpin structure exhibits all the structural elements implicated in termination of transcription in, prokaryotes.

Base Sequence↗

Molecular cloning, sequencing and sequence analysis of the fox-2 gene of Neurospora crassa encoding the multifunctional beta-oxidation protein.

We present the molecular cloning and sequencing of genomic and cDNA clones of the fox-2 gene of Neurospora crassa, encoding the multifunctional beta-oxidation protein (MFP). The coding region of the fox-2 gene is interrupted by three introns, one of which appears to be inefficiently spliced out. The encoded protein comprises 894 amino acid residues and exhibits 45% and 47% sequence identity with the MFPs of Candida tropicalis and Saccharomyces cerevisiae, respectively. Sequence analysis identifies three regions of the fungal MFPs that are highly conserved. These regions are separated by two segments that resemble linkers between domains of other MFPs, suggesting a three-domain structure. The first and second conserved regions of each MFP are homologous to each other and to members of the short-chain alcohol dehydrogenase family. We discuss these homologies in view of recent findings that fungal MFPs contain enoyl-CoA hydratase 2 and D-3-hydroxyacyl-CoA dehydrogenase activities, converting trans-2-enoyl-CoA via D-3-hydroxyacyl-CoA to 3-ketoacyl-CoA. In contrast to its counterparts in yeasts, the Neurospora MFP does not have a C-terminal sequence resembling the SKL motif involved in protein targeting to microbodies.

3-Hydroxyacyl CoA Dehydrogenases↗

Complete nucleotide sequence of the 6 kb element and conserved cytochrome b gene sequences among Indian isolates of Plasmodium falciparum.

The malaria parasite contains a nuclear genome with 14 chromosomes and two extrachromosomal DNA molecules of 6 kb and 35 kb in size. The smallest genome, known as the 6 kb element or mitochondrial DNA, has been sequenced from several Plasmodium falciparum isolates because this is a potential drug target. Here we describe the complete nucleotide sequence of this element from an Indian isolate of P. falciparum. It is 5967 bp in size and shows 99.6% homology with the 6 kb element of other isolates. The element contains three open reading frames for mitochondrial proteins-cytochrome oxidase subunit I (CoI), subunit III (CoIII) and cytochrome b (Cyb) which were found to be expressed during blood stages of the parasite. We have also sequenced the entire cyb gene from several Indian isolates of P. falciparum. The rate of mutation in this gene was very low since 12 of 14 isolates showed the identical sequence. Only one isolate showed a maximum change in five amino acids whereas the other isolate showed only one amino acid change. However, none of the Indian isolates showed any change in those amino acids of cyb which are associated with resistance to various drugs as these drugs are not yet commonly used in India.

Amino Acid Sequence↗

Mammalian mitochondrial ribosomal proteins. N-terminal amino acid sequencing, characterization, and identification of corresponding gene sequences.

The integrity of healthy mitochondria is supposed to depend largely on proper mitochondrial protein biosynthesis. Mitochondrial ribosomal proteins (MRPs) are directly involved in this process. To identify mammalian mitochondrial ribosomal proteins and their corresponding genes, we purified mature rat MRPs and determined 12 different N-terminal amino acid sequences. Using this peptide information, data banks were screened for corresponding DNA sequences to identify the genes or to establish consensus cDNAs and to characterize the deduced MRP open reading frames. Eight different groups of corresponding mammalian MRPs constituted from human, mouse, and rat origin were identified. Five of them show significant sequence similarities to bacterial and/or yeast mitochondrial ribosomal proteins. However, MRPs are much less conserved in respect to the amino acid sequence among species than cytoplasmic ribosomal proteins of eukaryotes and bacteria.

Amino Acid Sequence↗

Integrated graphical analysis of protein sequence features predicted from sequence composition.

Several protein sequence analysis algorithms are based on properties of amino acid composition and repetitiveness. These include methods for prediction of secondary structure elements, coiled-coils, transmembrane segments or signal peptides, and for assignment of low-complexity, nonglobular, or intrinsically unstructured regions. The quality of such analyses can be greatly enhanced by graphical software tools that present predicted sequence features together in context and allow judgment to be focused simultaneously on several different types of supporting information. For these purposes, we describe the SFINX package, which allows many different sets of segmental or continuous-curve sequence feature data, generated by individual external programs, to be viewed in combination alongside a sequence dot-plot or a multiple alignment of database matches. The implementation is currently based on extensions to the graphical viewers Dotter and Blixem and scripts that convert data from external programs to a simple generic data definition format called SFS. We describe applications in which dot-plots and flanking database matches provide valuable contextual information for analyses based on compositional and repetitive sequence features. The system is also useful for comparing results from algorithms run with a range of parameters to determine appropriate values for defaults or cutoffs for large-scale genomic analyses.

Amino Acid Motifs↗

Probabilistic alignment detects remote homology in a pair of protein sequences without homologous sequence information.

Dynamic programming (DP) and its heuristic algorithms are the most fundamental methods for similarity searches of amino acid sequences. Their detection power has been improved by including supplemental information, such as homologous sequences in the profile method. Here, we describe a method, probabilistic alignment (PA), that gives improved detection power, but similarly to the original DP, uses only a pair of amino acid sequences. Receiver operating characteristic (ROC) analysis demonstrated that the PA method is far superior to BLAST, and that its sensitivity and selectivity approach to those of PSI-BLAST. Particularly for orphan proteins having few homologues in the database, PA exhibits much better performance than PSI-BLAST. On the basis of this observation, we applied the PA method to a homology search of two orphan proteins, Latexin and Resuscitation-promoting factor domain. Their molecular functions have been described based on structural similarities, but sequence homologues have not been identified by PSI-BLAST. PA successfully detected sequence homologues for the two proteins and confirmed that the observed structural similarities are the result of an evolutional relationship.

Amino Acids↗

Two sequence-ready contigs spanning the two copies of a 200-kb duplication on human 21q: partial sequence and polymorphisms.

Physical mapping across a duplication can be a tour de force if the region is larger than the size of a bacterial clone. This was the case of the 170- to 275-kb duplication present on the long arm of chromosome 21 in normal human at 21q11.1 (proximal region) and at 21q22.1 (distal region), which we described previously. We have constructed sequence-ready contigs of the two copies of the duplication of which all the clones are genuine representatives of one copy or the other. This required the identification of four duplicon polymorphisms that are copy-specific and nonallelic variations in the sequence of the STSs. Thirteen STSs were mapped inside the duplicated region and 5 outside but close to the boundaries. Among these STSs 10 were end clones from YACs, PACs, or cosmids, and the average interval between two markers in the duplicated region was 16 kb. Eight PACs and cosmids showing minimal overlaps were selected in both copies of the duplication. Comparative sequence analysis along the duplication showed three single-basepair changes between the two copies over 659 bp sequenced (4 STSs), suggesting that the duplication is recent (less than 4 mya). Two CpG islands were located in the duplication, but no genes were identified after a 36-kb cosmid from the proximal copy of the duplication was sequenced. The homology of this chromosome 21 duplicated region with the pericentromeric regions of chromosomes 13, 2, and 18 suggests that the mechanism involved is probably similar to pericentromeric-directed mechanisms described in interchromosomal duplications.

Cell Line↗

DNA sequence analysis of spontaneous histidine mutations in a polA1 strain of Escherichia coli K12 suggests a specific role of the GTGG sequence.

Spontaneously arising histidine mutations in an Escherichia coli K12 strain deficient for DNA polymerase I were analysed at the DNA sequence level. We screened approximately 150,000 colonies and isolated 106 histidine auxotrophs. Of these, 98 were unstable hisC mutations; 12 representative mutants analysed were shown to have arisen by the excision of a single quadruplet repeat in the sequence 5'-GCTGGCTGGCTGGCTG-3'. Of the eight mutations at other sites, three hisA deletions and one hisD deletion occurred as a consequence of misalignment of tandemly repeated pentamers (hisD) or decamers (hisA). A single hisA point mutation was found to be a missense mutation. Two extended deletions, covering the his operon were not analysed. We could not identify the hisC deletion by sequencing. We conclude that polA1 is a strong mutator that induces mutations mostly of the minus frameshift and deletion type by a Streisinger-type of mispairing in repetitive DNA sequences. Finally, the possible role of a 5'-GTGG-3' sequence and its inverted or direct complements, which are found in the vicinity of all the deletions and frameshifts, is discussed.

Base Sequence↗

The nucleotide sequence of the sigma factor gene ntrA (rpoN) of Azotobacter vinelandii: analysis of conserved sequences in NtrA proteins.

The nucleotide sequence of the Azotobacter vinelandii ntrA gene has been determined. It encodes a 56916 Dalton acidic polypeptide (AvNtrA) with substantial homology to NtrA from Klebsiella pneumoniae (KpNtrA) and Rhizobium meliloti (RmNtrA). NtrA has been shown to act as a novel RNA polymerase sigma factor but the predicted sequence of AvNtrA substantiates our previous analysis of KpNtrA in showing no substantial homology to other known sigma factors. Alignment of the predicted amino acid sequences of AvNtrA, KpNtrA and RmNtrA identified three regions; two showing greater than 50% homology and an intervening sequence of less than 10% homology. The predicted protein contains a short sequence near the centre with homology to a conserved region in other sigma factors. The C-terminal region contains a region of homology to the beta' subunit of RNA polymerase (RpoC) and two highly conserved regions one of which is significantly homologous to known DNA-binding motifs. In A. vinelandii, ntrA is followed by another open reading frame (ORF) which is highly homologous to a comparable ORF downstream of ntrA in K. pneumoniae and R. meliloti.

Amino Acid Sequence↗

Examination of protein sequence homologies: V. New perspectives on evolution between bacterial and chloroplast-type ferredoxins inferred from sequence evidence.

Sequence homologies among 34 chloroplast-type ferredoxins were examined using a computer program that quantitatively evaluates the extent of sequence similarity as a correlation coefficient. The resultant alignment contains six gaps representing insertions or deletions of some residues, all of which are located such that they precisely preserve the domains of structural fragments as determined by crystallographic data on Spirulina platensis ferredoxin. In the search for any total correlation between the chloroplast-type and 27 bacterial ferredoxins, 1891 comparison matrices prepared for possible combinations indicated that the bacterial basal sequence of 55 residues has been conserved evolutionarily in the chloroplast-type sequences corresponding to residue positions 36-90 of Spirulina platensis ferredoxin. In addition, the bacterial "connector sequence" region was found to be conserved. These findings strongly suggest that the bacterial and chloroplast-type ferredoxins descended from a common ancestor, and branched off after the bacterial gene duplication, whereas the chloroplast-type ferredoxins originally were generated by duplicating the already duplicated bacterial gene, i.e., by "double-duplication."

Amino Acid Sequence↗

A highly conserved sequence in H1 histone genes as an oligonucleotide hybridization probe: isolation and sequence of a duck H1 gene.

A 3.5-kb HindIII fragment of a histone gene cluster was isolated from a recombinant phage out of a duck genomic library. This DNA contains a duck H1 gene and its flanking sequences. The hybridization probe, which was used to screen for the H1 gene, had been designed on the basis of a comparative analysis of available H1 gene and protein data. Most H1 histones contain repeated motifs in their C-terminal domain, and these form part of an octapeptide (ser pro lys lys ala lys lys pro) that is highly conserved in many H1 histone proteins. A comparison of the duck H1 described here with two different published chicken H1 histone sequences reveals conservative amino acid exchanges at 22 (of 217 and 218, respectively) positions. The homology is maintained at the flanking sequences, and includes the putative H1 histone gene-specific signal structures and the established 3' stem and loop structures and the CAAGA box. The duck H1 gene and its flanking sequence have been found in identical arrangements in two recombinant bacteriophages, but minor sequence variations and genomic Southern blotting after HindIII digestion suggest that we have either isolated alleles of this genome segment or that the gene described may occur twice per haploid duck genome.

Amino Acid Sequence↗

Random sheared fosmid library as a new genomic tool to accelerate complete finishing of rice (Oryza sativa spp. Nipponbare) genome sequence: sequencing of gap-specific fosmid clones uncovers new euchromatic portions of the genome.

The International Rice Genome Sequencing Project has recently announced the high-quality finished sequence that covers nearly 95% of the japonica rice genome representing 370 Mbp. Nevertheless, the current physical map of japonica rice contains 62 physical gaps corresponding to approximately 5% of the genome, that have not been identified/represented in the comprehensive array of publicly available BAC, PAC and other genomic library resources. Without finishing these gaps, it is impossible to identify the complete complement of genes encoded by rice genome and will also leave us ignorant of some 5% of the genome and its unknown functions. In this article, we report the construction and characterization of a tenfold redundant, 40 kbp insert fosmid library generated by random mechanical shearing. We demonstrated its utility in refining the physical map of rice by identifying and in silico mapping 22 gap-specific fosmid clones with particular emphasis on chromosomes 1, 2, 6, 7, 8, 9 and 10. Further sequencing of 12 of the gap-specific fosmid clones uncovered unique rice genome sequence that was not previously reported in the finished IRGSP sequence and emphasizes the need to complete finishing of the rice genome.

Base Sequence↗

Reverse transcriptase domain sequences from Mungbean (Vigna radiata) LTR retrotransposons: sequence characterization and phylogenetic analysis.

The conserved domains of reverse transcriptase (RT) genes of Ty1-copia and Ty3-gypsy groups of long terminal repeat (LTR) retrotransposons were amplified from mungbean (Vigna radiata) genome using degenerate primers, cloned and sequenced. Among these 34% and 65% of respective clones of copia and gypsy RT sequences possessed stop codons or frame-shifts or both. The RT sequences corresponding to both the groups exhibit significant levels of heterogeneity. Presence of mungbean copia and gypsy RT sequences in other papilionoid legumes of the same (Phaseoleae) and different lineages (Loteae, Trifoleae, Cicereae) indicates existence of these elements prior to the radiation of papilionoid legumes and also supports the recent interpretations of close relationship between Phaseoleae and Loteae tribes of Papilionoideae subfamily. On the other hand significant homologies of some mungbean copia as well as gypsy RT sequences with those of unrelated plant species suggest their origin from different plant lineages and also that heterogeneous population of related elements were already existed throughout (even before the divergence of monocot and dicot) the evolution of these genera from their common ancestor.

Amino Acid Sequence↗