Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The nucleotide sequence of the gene encoding the Newcastle disease virus membrane protein and comparisons of membrane protein sequences.

The nucleotide sequence of a cloned cDNA copy of the mRNA encoding the Newcastle disease virus (NDV) membrane (M) protein was determined. A single open reading frame in the sequence encodes a protein of 364 amino acids with a calculated mol wt of 39,742. The predicted protein sequence does not contain extensive hydrophobic regions. The sequence does contain eight pairs of basic amino acid residues, five of which are located in the carboxyl-terminal half of the sequence. Comparisons of the NDV M protein sequence with other paramyxovirus M protein sequences reveals little homology common to all sequences. There is only 17% homology with Sendai virus M protein. A short region of homology with the VSV M protein sequence, was, however, found.

Amino Acid Sequence↗

Nucleotide sequence of the gene encoding the Newcastle disease virus hemagglutinin-neuraminidase protein and comparisons of paramyxovirus hemagglutinin-neuraminidase protein sequences.

The nucleotide sequence of cloned cDNA copies of the mRNA encoding the Newcastle disease virus (NDV), strain A-V, hemagglutinin-neuraminidase (HN) protein was determined. A single open reading frame in the sequence encodes a protein of 570 amino acids with a calculated molecular weight of 62,280. The predicted protein sequence contains only one obvious potential membrane spanning region, located 27 amino acids from the amino terminus of the sequence. The predicted sequence contains 6 glycosylation sites and 14 cysteine residues. Comparison of the NDV HN protein sequence with three other paramyxovirus HN protein sequences reveals two regions that have homologies in all four sequences. The conserved cysteine residues are clustered in these two regions. One conserved region is located near the middle of the predicted sequence while the second region is in the carboxy terminal third of the molecule. The presence of conserved regions suggests the importance of these areas of the molecule in the structure or function of the protein.

Amino Acid Sequence↗

Amino acid sequence of human histidine-rich glycoprotein derived from the nucleotide sequence of its cDNA.

A lambda gt 11 library containing cDNA inserts prepared from human liver mRNA has been screened with an affinity-purified antibody to human histidine-rich glycoprotein (HRG) and then with a restriction fragment isolated from the 5' end of the largest cDNA insert obtained by antibody screening. A number of positive clones were identified and shown to code for HRG by DNA sequence analysis. A total of 2067 nucleotides were determined by sequencing 3 overlapping cDNA clones, which included 121 nucleotides of 5'-noncoding sequence, 54 nucleotides coding for a leader sequence of 18 amino acids, 1521 nucleotides coding for the mature protein of 507 amino acids, a stop codon of TAA, and 352 nucleotides of 3'-noncoding sequence followed by a poly(A) tail of 16 nucleotides. The length of the noncoding sequence of the 3' end differed in several clones, but each contained a polyadenylylation or processing sequence of AATAAA followed by a poly(A) tail. More than half of the amino acid sequence of HRG consisted of five different types of internal repeats. Within the last 3 internal repeats (type V), there were 12 tandem repetitions of a 5 amino acid segment with a consensus sequence of Gly-His-His-Pro-His. This repeated portion, referred to as a "histidine-rich region", contained 53% histidine and showed a high degree of similarity to a histidine-rich region of high molecular weight kininogen.

Amino Acid Sequence↗

Data mining for simple sequence repeats in expressed sequence tags from barley, maize, rice, sorghum and wheat.

Plant genomics projects involving model species and many agriculturally important crops are resulting in a rapidly increasing database of genomic and expressed DNA sequences. The publicly available collection of expressed sequence tags (ESTs) from several grass species can be used in the analysis of both structural and functional relationships in these genomes. We analyzed over 260000 EST sequences from five different cereals for their potential use in developing simple sequence repeat (SSR) markers. The frequency of SSR-containing ESTs (SSR-ESTs) in this collection varied from 1.5% for maize to 4.7% for rice. In addition, we identified several ESTs that are related to the SSR-ESTs by BLAST analysis. The SSR-ESTs and the related sequences were clustered within each species in order to reduce the redundancy and to produce a longer consensus sequence. The consensus and singleton sequences from each species were pooled and clustered to identify cross-species matches. Overall a reduction in the redundancy by 85% was observed when the resulting consensus and singleton sequences (3569) were compared to the total number of SSR-EST and related sequences analyzed (24 606). This information can be useful for the development of SSR markers that can amplify across the grass genera for comparative mapping and genetics. Functional analysis may reveal their role in plant metabolism and gene evolution.

Computational Biology↗

Putative in silico mapping of DNA sequences to livestock genome maps using SSLP flanking sequences.

In this study, an in silico approach was developed to identify homologies existing between livestock microsatellite flanking sequences and GenBank nucleotide sequences. Initially, 1955 bovine, 1570 porcine and 1121 chicken microsatellites were downloaded and the flanking sequences were compared with the nr and dbEST databases of GenBank. A total of 74 bovine, 44 porcine and 37 chicken microsatellite flanking sequences passed our criteria and had at least one significant match to human genomic sequence, genes/expressed sequence tags (ESTs) or both. GenBank annotation and BLAT searches of the UCSC human genome assembly revealed that 38 bovine, 13 porcine and 17 chicken microsatellite flanking sequences were highly similar to known human genes. Map locations were available for 67 bovine, 44 porcine and 21 chicken microsatellite flanking sequences, providing useful links in the comparative maps of humans and livestock. In support of our approach, 112 alignments with both microsatellite and match mapping information were located in the expected chromosomal regions based on previously reported syntenic relationships. The development of this in silico mapping approach has significantly increased the number of genes and EST sequences anchored to the bovine, porcine and chicken genome maps and the number of links in various human-livestock comparative maps.

Animals↗

Nucleotide sequence of the simian sarcoma virus genome: demonstration that its acquired cellular sequences encode the transforming gene product p28sis.

The complete nucleotide sequence of the proviral genome of simian sarcoma virus (SSV), an acute transforming retrovirus of primate origin, has been determined. Like other transforming viruses, SSV contains sequences derived from its helper virus, simian sarcoma-associated virus (SSAV), and a cell-derived (v-sis) insertion sequence. By comparison with the sequence of Moloney murine leukemia virus, it was possible to precisely localize and define sequences contributed by SSAV during the generation of SSV. Comparative sequence analysis of SSV and SSAV showed that SSAV provides regulatory sequences for initiation and termination of transcription of the SSV transforming gene. Moreover, coding sequences for the putative protein product of this gene appear to initiate from the amino terminus of the SSAV env gene. Antibodies to synthetic peptides derived from the carboxy and amino termini of the putative protein predicted by the open reading frame identified within v-sis specifically detect a Mr 28,000 protein, p28sis, in SSV-transformed cells. These and other findings confirm the predicted amino acid sequence of this protein and localize it to the coding region of the SSV transforming gene.

Amino Acid Sequence↗

Amino acid sequence of the alpha subunit of transducin deduced from the cDNA sequence.

Transducin, a GTP-binding protein involved in phototransduction in the vertebrate retina, belongs to a family of homologous coupling proteins that also includes Gs and Gi, the regulatory proteins of adenylate cyclase. Here we report the cDNA sequence and deduced amino acid sequence of transducin's alpha subunit (T alpha). The cDNA was isolated, by screening with an antibody probe, from a bovine retinal cDNA library in the expression vector lambda gt11. The 2.2-kilobase cDNA insert hybridized to a single 2.6-kilobase poly(A)+ RNA species present in extracts of bovine retina but not of bovine heart, liver, or brain. The nucleotide sequence of the cDNA revealed an open reading frame long enough to encode the entire 39-kDa T alpha polypeptide. The polypeptide sequence deduced from the cDNA would be composed of 350 amino acids and have a molecular weight of 39,971. Portions of the sequence matched reported amino acid sequences of T alpha tryptic fragments, including sites specifically ADP-ribosylated by cholera and pertussis toxins. The predicted sequence also includes four segments, ranging from 11 to 19 residues in length, that exhibit significant homology to sequences of GTP-binding proteins, including the ras proteins of man and yeast and the elongation factors of ribosomal protein synthesis in bacteria, EF-G and EF-Tu. In combination with previous functional studies of tryptic fragments of T alpha, the deduced amino acid sequence makes it possible to predict which portions of the polypeptide interact with other molecules involved in retinal phototransduction.

Amino Acid Sequence↗

Amino acid sequence of rabbit fast-twitch skeletal muscle calsequestrin deduced from cDNA and peptide sequencing.

Partial amino acid sequence analysis of rabbit fast-twitch skeletal muscle calsequestrin permitted the construction of synthetic oligonucleotides that were used as both primers and probes for the synthesis and isolation of cDNAs encoding calsequestrin from neonatal rabbit skeletal muscle libraries. The cDNA sequence encodes a processed protein of 367 residues with a Mr of 42,435 and a 28-residue amino-terminal signal sequence. The deduced amino acid sequence agreed closely with the portions of the mature protein that were sequenced using standard protein sequencing. The neonatal protein, however, contains an acidic carboxyl-terminal extension not present in the adult protein, suggesting that the cDNA sequence may have arisen from an alternatively spliced neonatal transcript. A single transcript of 1.9-2.0 kilobases was seen in neonatal skeletal muscle mRNA. A glycosylation site and two potential phosphorylation sites were detected. Although the protein contains about two acidic residues for each Ca2+ bound, there is no repeating distribution of acidic residues and no evidence of EF hand structures. Hydropathy plots show no transmembrane sequences, and structural analyses suggest that less than half of the protein is likely to be highly structured. This sequence defines the characteristics of a class of high-capacity, moderate-affinity, Ca2+ binding proteins.

Amino Acid Sequence↗

Shotgun sequencing of the human transcriptome with ORF expressed sequence tags.

Theoretical considerations predict that amplification of expressed gene transcripts by reverse transcription-PCR using arbitrarily chosen primers will result in the preferential amplification of the central portion of the transcript. Systematic, high-throughput sequencing of such products would result in an expressed sequence tag (EST) database consisting of central, generally coding regions of expressed genes. Such a database would add significant value to existing public EST databases, which consist mostly of sequences derived from the extremities of cDNAs, and facilitate the construction of contigs of transcript sequences. We tested our predictions, creating a database of 10,000 sequences from human breast tumors. The data confirmed the central distribution of the sequences, the significant normalization of the sequence population, the frequent extension of contigs composed of existing human ESTs, and the identification of a series of potentially important homologues of known genes. This approach should make a significant contribution to the early identification of important human genes, the deciphering of the draft human genome sequence currently being compiled, and the shotgun sequencing of the human transcriptome.

Animals↗

Complete nucleotide sequence and deduced polypeptide sequence of a nonmuscle myosin heavy chain gene from Acanthamoeba: evidence of a hinge in the rodlike tail.

We have completely sequenced a gene encoding the heavy chain of myosin II, a nonmuscle myosin from the soil ameba Acanthamoeba castellanii. The gene spans 6 kb, is split by three small introns, and encodes a 1,509-residue heavy chain polypeptide. The positions of the three introns are largely conserved relative to characterized vertebrate and invertebrate muscle myosin genes. The deduced myosin II globular head amino acid sequence shows a high degree of similarity with the globular head sequences of the rat embryonic skeletal muscle and nematode unc 54 muscle myosins. By contrast, there is no unique way to align the deduced myosin II rod amino acid sequence with the rod sequence of these muscle myosins. Nevertheless, the periodicities of hydrophobic and charged residues in the myosin II rod sequence, which dictate the coiled-coil structure of the rod and its associations within the myosin filament, are very similar to those of the muscle myosins. We conclude that this ameba nonmuscle myosin shares with the muscle myosins of vertebrates and invertebrates an ancestral heavy chain gene. The low level of direct sequence similarity between the rod sequences of myosin II and muscle myosins probably reflects a general tolerance for residue changes in the rod domain (as long as the periodicities of hydrophobic and charged residues are largely maintained), the relative evolutionary "ages" of these myosins, and specific differences between the filament properties of myosin II and muscle myosins. Finally, sequence analysis and electron microscopy reveal the presence within the myosin II rodlike tail of a well-defined hinge region where sharp bending can occur. We speculate that this hinge may play a key role in mediating the effect of heavy chain phosphorylation on enzymatic activity.

Amino Acid Sequence↗

Comparison of genomic fragment and clone sequences within a long interspersed repeated sequence of the mouse genome.

The 393bp nucleotide sequence of a HindIII genomic fragment mapping within the major long interspersed repeated sequence family (MIF-1, Bam, L1) of mouse is reported and compared to clone sequences of the same region of this repeated sequence. The consensus of the clone sequences significantly differs from the genomic fragment sequence by additions and deletions that are inconsistent with the physical and biochemical properties of the genomic fragment. While alternative explanations could account for some of these differences, several aspects of the experimental results imply that cloning artifacts contribute to the discrepancies. Despite the differences between the clone and genomic fragment sequences, the biologically interesting features previously noted in clone sequences (promoter-like signals and an open reading frame) are conserved in the genomic fragment sequence.

Animals↗

Compilation of tRNA sequences and sequences of tRNA genes.

Maintained at the Universitat Bayreuth, Bayreuth, Germany, the Compilation of tRNA Sequences and Sequences of tRNA Genes is accessible at the URL http://www.tRNA.uni-bayreuth.de with mirror site located at the Institute of Protein Research, Pushchino, Russia (http://alpha.protres.ru/trnadbase). The compilation is a searchable, periodically updated database of currently available tRNA sequences. The present version of the database contains a new Genomic tRNA Compilation including the sequences of tRNA genes from genomic sequences published up to July 2003. It consists of about 5800 tRNA gene sequences from 111 organisms covering archaea, bacteria, higher and lower eukarya. The former Compilation of tRNA Genes (up to the end of 1998) and the updated Compilation tRNA Sequences (561 entries) are also supported by the new software. The database can be explored by using multiple search criteria and sequence templates. The database provides a service that allows to obtain statistical information on the occurrences of certain bases at given positions of the tRNA sequences. This allows phylogenic studies and search for identity elements in respect to interactions of tRNAs with various enzymes.

Animals↗

Molecular cloning and characterization of six genes, determination of gene order and intergenic sequences and leader sequence of mumps virus.

mRNA isolated from mumps virus-infected Vero cells was converted into cDNA and cloned into the PstI site of the plasmid pBR322. After screening with 32P-labelled cDNA synthesized from poly(A)+ RNA of uninfected or mumps virus-infected Vero cells, five different groups of virus-specific clones were obtained. The virus specificity of the clones was confirmed by Northern blot analysis, in which the cDNA inserts from the five different groups hybridized to mRNAs of about 2100, 1500, 1450, 2000 and 2200 nucleotides. By the use of oligonucleotides synthesized on the basis of sequences obtained from the five cDNA clones and mRNAs, the sequence of the intergenic and surrounding areas was determined. During genome sequencing, a separate gene was identified between the fusion protein (F) gene and the haemagglutinin-neuraminidase protein (HN) gene. Using oligonucleotides synthesized on the basis of the new gene sequence, cDNA clones with poly(A) were isolated from the cDNA library. The gene order was determined to be 3' NC-P-M-F-SH-HN-L 5' (where NC, P, M, SH, and L represent the genes for the nucleocapsid, phosphoprotein or polymerase-associated, matrix or membrane, small hydrophobic and large proteins respectively). There is one nucleotide between the P and M (A), M and F (A), and HN and L genes (G), two between the NC and P (AA) and SH and HN (3'-CG) genes, and seven between the F and SH genes (3' GAUUUUA) as intergenic sequence. The leader sequence at the 3' end of the genome has been determined by sequencing the dicistronic leader-NC mRNA using oligonucleotide primers. The sequence from the 3' terminus to the NC gene start of the mumps virus genome is similar in length (55 nucleotides) to that present in Sendai virus, Newcastle disease virus and parainfluenza virus type 3, and the first five nucleotides are conserved in all negative-stranded RNA virus genomes sequenced to date.

Base Sequence↗

Defining the sequence specificity of DNA-binding proteins by selecting binding sites from random-sequence oligonucleotides: analysis of yeast GCN4 protein.

We describe a new method for accurately defining the sequence recognition properties of DNA-binding proteins by selecting high-affinity binding sites from random-sequence DNA. The yeast transcriptional activator protein GCN4 was coupled to a Sepharose column, and binding sites were isolated by passing short, random-sequence oligonucleotides over the column and eluting them with increasing salt concentrations. Of 43 specifically bound oligonucleotides, 40 contained the symmetric sequence TGA(C/G)TCA, whereas the other 3 contained sequences matching six of these seven bases. The extreme preference for this 7-base-pair sequence suggests that each position directly contacts GCN4. The three nucleotide positions on each side of this core heptanucleotide also showed sequence preferences, indicating their effect on GCN4 binding. Interestingly, deviations in the core and a stronger sequence preference in the flanking region were found on one side of the central C . G base pair. Although GCN4 binds as a dimer, this asymmetry supports a model in which interactions on each side of the binding site are not equivalent. The random selection method should prove generally useful for defining the specificities of other DNA-binding proteins and for identifying putative target sequences from genomic DNA.

Base Sequence↗

Detection of inter-spread repeat sequence in genomic DNA sequence.

Various types of periodic patterns in nucleotide sequences are known to be very abundant in a genomic DNA sequence, and to play important biological roles such as gene expression, genome structural stabilization, and recombination. We present a new method, named "STEPSTONE", to find a specific periodic pattern of repeat sequence, inter-spread repeat, in which the tandem repeats of the conserved and the not-conserved regions appear periodically. In our method, at first, the data on periods of short repeat sequences found in a target sequence are stored as a hash data, and then are selected by application of an auto-correlation test in time series analysis. Among the statistically selected sequences, the inter-spread repeats are obtained by usual alignment procedures through two steps. To test the performance of our method, we examined the inter-spread repeats in Mycobacterium tuberculosis and Zamia paucijuga genomic sequences. As a result, our method exactly detected the repeats in the two sequences, being useful for identifying systematically the inter-spread repeats in DNA sequence.

Algorithms↗

cDNA sequence of a novel sex-limited protein (Slp) from mice constitutive for Slp expression. Sequence comparisons suggest that Slp has no functional role.

Murine sex-limited protein (Slp) is a serum protein that shares 95% sequence identity with murine complement component C4 but does not have C4 activity. Mouse strain B10.WR, which carries the H-2w7 haplotype, has up to 4 Slp genes and is unusual in that both males and females express Slp. Here we report the sequence of a complete pro-Slp cDNA from this strain that we designate Slpw7.2. We find that the Slpw7.2 sequence differs at multiple dispersed sites from three previously reported Slp sequences: two complete pro-Slp sequences, Slpw7.1 and SlpFM, from the B10.WR and FM strains, respectively, and a partial sequence from the B10.WR strain that is distinct from Slpw7.1 as well. A detailed comparison of the complete Slpw7.1, Slpw7.2, and SlpFM sequences reveals that nucleotide changes that alter the amino acid sequence (replacement substitutions) are accumulating at the same relative rate as changes that do not affect the amino acid sequence (silent substitutions); in addition, the amino acid changes themselves tend to be nonconservative. Our results suggest that at least three Slp genes are transcriptionally active in B10.WR mice; that the protein product of the Slpw7.2 transcript predominates in B10.WR serum; and that the Slp protein probably has no function. The Slp system may be particularly suitable for the study of the evolution in the absence of selective pressures of a gene that encodes a stable protein.

Amino Acid Sequence↗

Nucleotide sequence of bovine prolactin messenger RNA. Evidence for sequence polymorphism.

Hybrid molecules containing DNA sequences complementary to bovine pituitary mRNA were constructed in the Pst I site of pBR322 by the dC . dG tailing technique. Recombinant plasmids containing bovine prolactin (bPRL) sequences were amplified in bacteria and identified by hybridization to purified [32P]bPRL cDNA sequences. Nucleotide sequence analysis was performed on the inserts from two of the positive clones. One clone, pBPRL72, contained a 982-base pair insert that included 67 nucleotides of the 5'-untranslated region, the complete coding region of the preprolactin protein (690 nucleotides), and the entire 3'-untranslated region (150 nucleotides) of bPRL mRNA. The nucleotide sequence analysis of clone pBPRL72 predicted the sequence of a 30-amino acid signal peptide and confirmed the published amino acid sequence of the protein with one exception. A comparison of the pBPRL72 cDNA sequence with a second bPRL clone, pBPRL4, revealed four silent nucleotide differences. Three of the base changes occurred in the third position of amino acid codons, and one occurred in the 3'-noncoding region. The sequence polymorphism suggests the existence of alleles or multiple loci for bPRL that do not alter the protein structure.

Amino Acid Sequence↗

The complete sequences of the galago and rabbit beta-globin locus control regions: extended sequence and functional conservation outside the cores of DNase hypersensitive sites.

The locus control region (LCR) of mammalian beta-globin genes covers at least 17 kb at the 5' end of the gene cluster and has been implicated in chromatin domain opening, enhancement, and insulation from neighboring sequences. Functional dissection of the LCR has defined the minimal cores for four of the five major DNase hypersensitive sites (HSs) that mark this regulatory region. To examine fully the patterns of conserved sequences in the mammalian homologs to the beta-globin LCR, we determined the complete DNA sequence of the galago beta-globin LCR and completed previously unsequenced regions of the rabbit LCR. Simultaneous alignment of these sequences with the human, goat, and mouse LCRs revealed conserved sequences (phylogenetic footprints) detected using three largely independent methods. The most highly conserved segments are found both within the HS cores and in some but not all regions flanking the cores. These results argue for an extended pattern of well-conserved sequences, many of which lie outside the minimal cores, and we show that a key sequence required for domain opening by the region including HS3 maps about 1 kb 5' to the minimal core. Differential phylogenetic footprints, containing sequences conserved in nonhuman mammals but not in humans, are found primarily around HS3, consistent with some species-specific differences in function that may be important for differences in hemoglobin switching during development.

Animals↗