Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

The 5' flanking sequences of the mouse P-cadherin gene. Homologies to 5' sequences of the E-cadherin gene and identification of a first 215 base-pair intron.

A genomic clone containing the 5' region of the mouse P-cadherin gene has been isolated from Balb/c mice. A major feature of this genomic sequence is the presence of a first intron (Il), 215 bp long, located 48 bp downstream of the translation start ATG codon. The presence of Il has been detected both in Balb/c and C57BL/6 mouse strains, being located at the same position with respect to coding sequences as in the mouse E-cadherin and chicken L-CAM genes. The transcription initiation site of the mouse P-cadherin gene has been located at about 68 nt from the ATG start codon, giving an estimation for the size of the first exon of the mouse P-cadherin gene of 116 bp. The sequence of the 5' upstream region of the P-cadherin gene presents structural similarities with the recently described 5' region of the mouse E-cadherin gene: absence of a TATA box, presence of a CAAT box at -65, two putative AP2-binding motifs, at -101 and +31, and a GC-rich region containing a potential SP1-binding element at -88. However, no sequence homologous to the palindromic sequence, E-pal, found on the E-cadherin promoter has been found in the 5' region of the P-cadherin gene. These results indicate that, in contrast to a previous report, the mouse E and P-cadherin genes exhibit a similar genomic organization both containing 15 introns and a similar size for the first two exons.

Amino Acid Sequence↗

Nucleotide sequence of the Pseudomonas aeruginosa insertion sequence IS222: another member of the IS3 family.

Sequence analysis of the Pseudomonas aeruginosa insertion sequence element IS222 revealed it to be 1234 bp in size with 23 bp imperfect terminal inverted repeats. Insertion caused a 5-bp duplication of the insertion site. Two ORFs were identified, one of which, ORFA, could encode a basic (pI 10.5) polypeptide with a mass of 11,709. This sequence bears strong homology to the putative ORFA product from the Shigella dysenteriae insertion sequence element IS911, which is a member of the IS3 family of insertion elements. As with other members of this group the nucleotide sequence contains a "frameshift window" (AAAAAAG; M. Chandler and O. Fayet (1993). Mol. Microbiol. 7, 497-503) at which ribosome slippage can result in a fusion protein (ORFAB).

Amino Acid Sequence↗

Complete sequence of Leishmania RNA virus 1-4 and identification of conserved sequences.

In order to understand the coding strategies and identify potential cis-acting sequences in Leishmania RNA virus 1 (LRV1), a complete cDNA sequence was obtained for LRV1-4 and compared to the sequence reported for LRV1-1. The results show that the 5' end of LRV1 is conserved at the nucleotide level while open reading frames (ORFs) 2 and 3 are conserved at the amino acid level. A simple translation initiation consensus sequence is conserved at the 5' end of ORF2 but absent from ORF3, consistent with a possibility that ORF3 is expressed as a gag-pol fusion protein. Comparison of secondary structure predictions obtained for both isolates identified nucleotide sequences capable of forming conserved stem-loops at the virus termini and in the putative frameshift region between ORF2 and ORF3. Although direct evidence is lacking, the appearance of compensatory nucleotide substitutions suggests that the structures may form in vivo. Possible functions for the conserved structures are discussed.

Amino Acid Sequence↗

Nucleotide sequence of the E2-peplomer protein gene and partial nucleotide sequence of the upstream polymerase gene of transmissible gas gastroenteritis virus (Miller strain).

The E2-peplomer protein gene of the virulent Miller strain of transmissible gastroenteritis virus (TGEV) was sequenced from cDNA clones and compared to the E2 gene sequence of the avirulent Purdue strain. Sequence comparisons indicate that most amino acid differences occur in the N-terminal half of the E2-peplomer which represents the most exposed region of the protein. In addition, analysis of an incompletely sequenced open reading frame (ORF) to the immediate 5' side of the E2 gene indicates extensive sequence homology with the infectious bronchitis virus (IBV) F2 gene which is thought to encode a RNA polymerase.

Amino Acid Sequence↗

A gene showing sequence similarity to pectin esterase is specifically expressed in developing pollen of Brassica napus. Sequences in its 5' flanking region are conserved in other pollen-specific promoters.

Differential screening of a Brassica napus genomic library led to the isolation of the clone named Bp 19 containing a gene which is highly expressed during microspore development. The accumulation of Bp19 mRNA starts in uninucleate microspores, increases during development reaching a peak in the late stages but declines considerably in mature pollen. The nucleotide sequence of the entire coding region and of extended portions of the 5' and 3' flanking regions was determined. Several homologous cDNA clones were also isolated and sequenced. The Bp 19 gene contains a single intron of 137 bp and gives origin to a mRNA of ca. 1.9 kb which codes for a polypeptide of 584 amino acids. Bp 19 protein has an estimated molecular weight of 63 kilodaltons and has a highly hydrophobic amino terminal region which shows features of a signal peptide. The carboxy half of the Bp 19 protein, starting at amino acid 269, has striking sequence similarity to the pectin esterases of tomato and of the plant pathogen Erwinia chrysanthemi. Four short domains are extremely well conserved in all the three proteins and therefore could represent catalytic sites responsible for enzyme activity. Comparison of the 5' flanking region of the Bp 19 gene with the sequence of other pollen-specific promoters revealed the presence of several conserved regions. These short promoter sequences could correspond to regulatory elements responsible for pollen-specific gene expression.

Amino Acid Sequence↗

Expression enhancement of the Tn5 neomycin-resistance gene by removal of upstream ATG sequences and its use for probing heterologous upstream activating sequences in yeast.

We have constructed a series of promoter or upstream activating sequence (UAS)-probe plasmids carrying the Tn5-derived neomycin resistance gene whose seven additional ATG codons in the 5'-untranslated region were completely or partially removed. When the deleted version of the neo sequence retaining only one additional ATG (NeoD) was expressed under the control of a TDH3 promoter whose UAS was deleted, the transformed cells were unable to grow at a low concentration of the antibiotic G418. In contrast with this, yeast cells expressing the NeoC sequence and having no additional ATG exhibited a high level of G418-resistance. Moreover, the UAS-probe system using NeoD has been successfully applied for the identification of several E. coli DNA sequences that clearly function as UASs in yeast cells. Two of these prokaryotic sequences with UAS activity were identified as a part of the coding region of the tgt and the hydG gene, respectively.

Amino Acid Sequence↗

DNA sequence of the metC gene and its flanking regions from Salmonella typhimurium LT2 and homology with the corresponding sequence of Escherichia coli.

The DNA sequence of the Salmonella typhimurium metC gene and its flanking regions was determined. The metC gene contains an open reading frame of 1185 nucleotides encoding a polypeptide of 395 amino acids with a predicted molecular weight of 42,874 daltons. S1 nuclease mapping experiments located the transcription start site of the metC gene. The nucleotide sequence and the deduced amino acid sequence for the metC genes of S. typhimurium and Escherichia coli were compared. Although there are 279 nucleotide replacements, most do not change the amino acid sequence. Nucleotide sequence analysis of the flanking regions of the S. typhimurium metC gene shows that there is an open reading frame upstream and an open reading frame downstream of the gene. The existence of the divergently transcribed upstream open reading frame (designated ORF1) was confirmed by the construction of an ORF1-lacZ fusion. The transcription start site of ORF1 was determined by S1 nuclease mapping.

Amino Acid Sequence↗

Nucleotide sequence of a pregnancy-specific beta 1 glycoprotein gene family member. Identification of a functional promoter region and several putative regulatory sequences.

The pregnancy-specific beta 1 glycoprotein (PSG) genes encode a group of heterogeneous proteins produced in large amounts by the human syncytiotrophoblast. Their expression seems to be regulated at the transcriptional level during normal pregnancy. In the present work, we isolated from a human placental library a 17 kb genomic fragment corresponding to a member of the PSG multigene family. DNA sequence analysis of 1190 nucleotides upstream of the translational start and of the first intron, revealed the presence of several putative regulatory sequences. In a transient chloramphenicol acetyltransferase expression assay, 5' flanking sequences within 123 nucleotides upstream to the first major transcription initiation site, functioned as a strong promoter in COS-7 cells. Meanwhile, sequences 5' further upstream had the ability to abolish this promoter activity. The sequence analyzed did not contain any obvious TATA-like boxes or G+C-rich regions, suggesting the existence of unique promoter elements implicated in transcription initiation and regulation of this PSG gene family member.

Amino Acid Sequence↗

Homology of the HSV-2 "a-sequence" to cellular sequences.

Bgl-II fragments of the genome of Herpes simplex virus type 2 (HSV-2) HG-52 were cloned into the vector p-Neo and were used to screen the complete HSV-2 genome for regions cross-hybridizing with the genome of HEL cells. Most extensive cross-hybridizing activity was observed with a 530 bp SstII subfragment of the viral BamHI G DNA-fragment (contained in Bgl II F), which spans the joint and the viral a-sequence. From a lambda-L47 library, a cellular 15 kb HindIII DNA fragment was subcloned in pBR 322 which contained a 1920 bp SstII subfragment having strong cross-hybridizing activity with the 530 bp Sst II fragment of HSV-2 BamHI G. Within this 1920 bp Sst II fragment the cross-hybridizing activity was confined to a 230 bp Bgl I/Hpa II subfragment. This 230 bp fragment (including the flanking sequences) was analyzed in comparison to the viral a-sequence. Sequence data revealed a (G + C) content of 66% in the cellular and 81% in the viral DNA fragment, which is mainly determined by an extremely (G + C) rich 16-fold direct repeat (DR2) at the 5'-end. The homology between both DNA-fragments varies between 56% and 79% within the L-S inversion region. Both sequences, furthermore, show homology to the human c-myc protooncogene.

Animals↗

The neurotoxin-like sequence of human immunodeficiency virus gp120: a comparison of sequence data from patients with and without neurological symptoms.

A region of the human immunodeficiency virus type 1 (HIV-1) envelope glycoprotein gp 120 has been claimed previously to be homologous to parts of snake venom neurotoxins and rabies virus glycoprotein ("the neurotoxic loop"). We have determined DNA sequences directly from a polymerase chain reaction amplified fragment corresponding to this region of HIV-1 gp 120 and have translated these to protein sequences. This was performed with the prototype HIVSF2 isolate and several Swedish HIV-1 strains, which were precultivated from blood cells or cerebrospinal fluid (CSF) or were directly obtained from CSF cells of patients with and without neurological symptoms. The results show that there are sequence similarities between a short segment of gp120 of clinical HIV-1 strains and the neurotoxic loop. The strains of patients with neurological symptoms did not, however, show a genetic shift of their sequences towards a greater similarity to the sequences of snake venom neurotoxins and rabies virus glycoprotein as compared to the strains of asymptomatic individuals.

AIDS Dementia Complex↗

The sequence upstream of the -10 consensus sequence modulates the strength and induction time of stationary-phase promoters in Escherichia coli.

We constructed a library of synthetic stationary-phase promoters for Escherichia coli. For designing the promoters, the known -10 consensus sequence, as well as the extended -10 region, and an A/T-rich region downstream of the -10 region were kept constant, whereas sequences from -37 to -14 were partially or completely randomised. For detection and selection of stationary-phase promoters, green fluorescent protein (GFP) with enhanced fluorescence was used. To establish the library, 33 promoters were selected, which differ in strength from 670 to more than 13,000 specific fluorescence units, indicating that the strength of promoters can be modulated by the sequence upstream of the -10 region. DNA sequencing revealed a preferential insertion of nucleotides depending on the position. By expressing the promoters in an rpoS-deficient strain, a special group of stationary-phase promoters was identified, which were expressed exclusively or preferentially by RNA polymerase holoenzyme Esigma(s). The DNA sequence of these promoters differed significantly in the region from -25 to -16. Furthermore, it was shown that the DNA curvature of the promoter region had no effect on promoter strength. The broad range of promoter activities make these promoters very suitable for fine-tuning of gene expression and for cost-effective large-scale applications in industrial bioprocesses.

Base Sequence↗

A New York isolate of Soil-borne wheat mosaic virus differs considerably from the Nebraska type strain in the nucleotide sequences of various coding regions but not in the deduced amino acid sequences.

A wheat-infecting furovirus found in Tompkins County, New York, U.S.A. was identified as a strain of Soil-borne wheat mosaic virus (SBWMV) by means of sequence analyses of portions of its RNA 1 and 2. The nucleotide sequences of several of its genes differed by c. 9 to 12% from those of the corresponding genome regions of the Nebraska type strain of SBWMV. The deduced amino acid sequences of the putative translation products, however, suggested much closer relationships. Thus, the amino acid sequences of the coat proteins of the two virus strains were 100% identical despite the fact that their coding regions differed in as many as 68 nucleotide positions. The New York (NY) strain of SBWMV is possibly closely related to an isolate from Illinois for which so far only the nucleotide sequences of its coat protein gene and the 5' untranslated region of its RNA 2 are known.

Amino Acid Sequence↗

Nucleotide sequence of the full-length mouse lamin C cDNA and its deduced amino-acid sequence.

We have cloned and sequenced the cDNA comprizing the entire coding region and several hundred base-pairs of its flanks for the mouse nuclear envelope protein lamin C mRNA. The nucleotide sequence and the deduced amino-acid sequence of the mouse lamin C are compared with the previously published human lamin A/C sequences with respect to (a) the general organisation, (b) homologies, (c) predictions for the essential structural characteristics of lamins and (d) the localization of the most conserved region. Moreover, the mouse lamin C sequence presented allows the first intraspecies comparison between A/C-type and B-type lamins.

Amino Acid Sequence↗

Sequence of human eosinophil-derived neurotoxin cDNA: identity of deduced amino acid sequence with human nonsecretory ribonucleases.

Several clones of human eosinophil-derived neurotoxin (EDN) cDNA have been isolated from a lambda gt10 cDNA library prepared from mRNA derived from noninduced HL-60 cells. The amino acid (aa) sequence deduced from the coding sequence of the EDN cDNA is identical to the aa sequence of urinary nonsecretory RNase. Comparison of the aa and/or nucleotide (nt) sequences of EDN and other proteins possessing ribonucleolytic activity, namely bovine seminal RNase, human and rat pancreatic RNases, eosinophil cationic protein (ECP), and human angiogenin, shows extensive identity at half-cystine residues and at aa of active sites. Differences in aa sequences at the active sites are often the result of single nt changes in the codons. The data presented here support the concept of a RNase gene superfamily containing secretory and nonsecretory RNases, angiogenin, EDN and ECP.

Amino Acid Sequence↗

Cloning and sequencing of 5' flanking sequence from the gene encoding 2S storage protein, from two Brassica species.

Using oligodeoxyribonucleotide primers and the polymerase chain reaction, we have cloned and sequenced about 1.2 kb of upstream sequences from two members of the 2S seed storage protein-encoding gene family from Brassica juncea and B. oleracea. The two sequences bear more than 90% homology and have characteristic seed-specific promoter motifs. The high degree of sequence conservation indicates that this napin-encoding gene family evolved earlier than the divergence of the three primary Brassica species and their amphidiploids, and the sequences have been conserved due to some metabolic constraints in seed development.

2S Albumins, Plant↗

Patterns of intra- and interarray sequence variation in alpha satellite from the human X chromosome: evidence for short-range homogenization of tandemly repeated DNA sequences.

A number of processes, such as sequence conversion, unequal crossingover, and molecular drive, have been postulated to explain the homogenization of tandemly repeated DNA families. To investigate the nature and extent of such processes in the alpha satellite family of centromeric DNA, we determined the nucleotide sequence of approximately 700 bp from each of 40 representative alpha satellite repeats from six sources of human X chromosomes, obtaining a total of approximately 28 kb of sequence data. Sequence divergence among the repeats examined was low, with an average pairwise difference of approximately 1%. Pairwise comparisons of all repeats indicate that the degree of similarity for those repeats in physical proximity (within approximately 15 kb) of each other is significantly greater than that for randomly located repeats, from either the same or different X chromosomes, suggesting that the mechanisms predicted to homogenize these arrays are effectively short-range in action. Analysis of individual patterns of sequence variation allows the assignment of haplotypes for five high-copy-number diagnostic positions and reveals distinct positions of equilibrium and disequilibrium within the repeat. These analyses address hypotheses about the origin of the observed patterns of variation throughout alpha satellite evolution.

Base Sequence↗

Comparative analysis of genome sequences of three isolates of Orf virus reveals unexpected sequence variation.

Orf virus (ORFV) is the type species of the Parapoxvirus genus. Here, we present the genomic sequence of the most well studied ORFV isolate, strain NZ2. The NZ2 genome is 138 kbp and contains 132 putative genes, 88 of which are present in all analyzed chordopoxviruses. Comparison of the NZ2 genome with the genomes of 2 other fully sequenced isolates of ORFV revealed that all 3 genomes carry each of the 132 genes, but there are substantial sequence variations between isolates in a significant number of genes, including 9 with inter-isolate amino acid sequence identity of only 38-79%. Each genome has an average of 64% G+C but each has a distinctive pattern of substantial deviation from the average within particular regions of the genome. The same pattern of variation was also seen in the genome of another parapoxvirus species and was clearly unlike the uniform patterns of G+C content seen in all other genera of chordopoxviruses. The availability of genomic sequences of three orf virus isolates allowed us to more accurately assess likely coding regions and thereby revise published data for 24 genes and to predict two previously unrecognized genes.

Base Composition↗

Completion of the DNA sequence of mouse adenovirus type 1: sequence of E2B, L1, and L2 (18-51 map units).

The DNA sequence of 9991 nt, corresponding to 18-51 map units of mouse adenovirus type 1 (MAV-1), was determined, completing the sequence of the Larsen strain of MAV-1. The length of the complete MAV-1 genome is 30,946 nucleotides, consistent with previous experimental estimates. The 18-51 map unit region encodes early region 2B proteins necessary for adenoviral replication as well as late region L1 and L2 structural and packaging proteins. Sequence comparison in this region with human adenoviruses indicates broad similarities, including colinear preservation of all recognized open reading frames (ORFs), with highest amino acid identity occurring in the DNA polymerase and polypeptide III (penton base subunit) ORFs. Virus-associated (VA) RNA is not encoded in the region where VA RNAs are found in the human adenoviruses, between E2B and L1, nor is it encoded anywhere in the entire MAV-1 genome. The MAV-1 polypeptide III lacks the arginine-glycine-aspartic acid (RGD) motif which is involved in an association with cell-surface integrins. Only one RGD sequence is found in an identified coding region in the entire MAV-1 genome. Similar to the porcine adenovirus, this RGD sequence is found in the C-terminus of the MAV-1 fiber protein.

Adenovirus E2 Proteins↗