Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Amino acid sequence of heavy chain from Xenopus laevis IgM deduced from cDNA sequence: implications for evolution of immunoglobulin domains.

Present understanding of the evolution of immunoglobulins is derived almost entirely from studies of a few mammalian species. To obtain information about immunoglobulin genes in Xenopus laevis, a cDNA library was prepared in the expression vector lambda gt11 from mitogen-stimulated splenocytes of this species. Of approximately equal to 50,000 clones screened, 18 were found to express IgM epitopes. One of these, lambda XIg14, hybridized with RNA of RNA of approximately equal to 2 kilobases from splenocytes. The insert of this clone appears to encode a variable region and part of a mu constant region; that of another clone, lambda XIg8, appears to encode a variable region and a complete mu constant region. Both inserts contain sequence corresponding to the three gene segments (VH, DH, and JH) that encode heavy-chain variable regions. The heavy-chain constant region (CH) encoded by lambda XIg8 has the characteristic features of C mu, including a four-domain structure and a carboxyl-terminal tail. The amino acid sequences of two mu-chain peptides agree with the cDNA sequence. The identity in amino acid sequence between the corresponding Xenopus and mouse C mu domains ranges from 31 to 47%. The C mu domains vary in the extent to which their sequences resemble the sequences of other immunoglobulins, consistent with previous suggestions that the immunoglobulin domains have an independent evolutionary history.

Amino Acid Sequence↗

Weighting in sequence space: a comparison of methods in terms of generalized sequences.

Four methods for weighting aligned biological sequences have recently appeared that differ mathematically, philosophically, and in their results. Thus, while there is consensus about the need to weight sequences, the method to use is contentious. A geometric analysis based on a continuous sequence space is presented that provides a common framework in which to compare the methods. It is concluded that there are two "best" methods. When the sequences are known to be phylogenetically related and a tree can be generated without introducing excessive stress into the data, the method of Altschul et al. [Altschul, S. F., Carroll, R. J. & Lipman, D. J. (1989) J. Mol. Biol. 207, 647-653] is appropriate. When the sequences are not known to be phylogenetically related or a tree cannot be produced without unduly distorting the distances between the sequences, a modification of the method of Sibbald and Argos [Sibbald, P. R. & Argos, P. (1990) J. Mol. Biol. 216, 813-818] is preferable.

Amino Acid Sequence↗

Simple sequence DNA associated with near sequence identity of the 3'-flanking regions of rat cytochrome P450b and P450e genes.

In rat liver, the two major phenobarbital (PB)-inducible cytochrome P450s, P450b (P450IIB1), and P450e (P450IIB2), are encoded by approximately 2.1-kb mRNAs showing more than 97% nucleotide sequence identity. Almost half of the sequence differences are concentrated in two short divergent segments in exon 7. An additional 4.8-kb P450b/P450e RNA, inducible by Aroclor and by PB, hybridizes with a classical P450b probe and with the 3' extension of the PB23 insert, a cloned P450b-like cDNA (Affolter et al., DNA 5, 209-218, 1986). The 4.8-kb form has now been detected in several rats under a variety of induction conditions. DNA sequencing of the 5' portion of the PB23 insert showed it is derived from a P450b gene. The 4.8-kb RNA hybridized with an oligonucleotide probe that recognizes P450b mRNA, but did not hybridize detectably with one that recognizes P450e mRNA. The 4.8-kb RNA and the PB23 insert doubtless represent P450b RNAs polyadenylated at a downstream site. DNA sequence analysis of the PB23 3' extension (which represents the 3' end of the P450b gene) and of the P450e gene 3'-flanking region demonstrated that the near identity of the P450b and P450e genes extends for at least 920 bp beyond the major polyadenylation site. This region was doubtless part of the original duplication that gave rise to the P450b and P450e genes. A short divergent segment is present in the 3'-flanking region of near sequence identity; the segment is embedded in simple sequence DNA that contains a mixed alternating pyrimidine/purine tract.

Animals↗

Prediction of the coding sequences of unidentified human genes. X. The complete sequences of 100 new cDNA clones from brain which can code for large proteins in vitro.

As an extension of our cDNA analysis for deducing the coding sequences of unidentified human genes, we have newly determined the sequences of 100 cDNA clones from a set of size-fractionated human brain cDNA libraries, and predicted the coding sequences of the corresponding genes, named KIAA0611 to KIAA0710. In vitro transcription-coupled translation assay was applied as the first screening to select cDNA clones which produce proteins with apparent molecular mass of 50 kDa and over. One hundred unidentified cDNA clones thus selected were then subjected to sequencing of entire inserts. The average size of the inserts and corresponding open reading frames was 4.9 kb and 2.8 kb (922 amino acid residues), respectively. Computer search of the sequences against the public databases indicated that predicted coding sequences of 87 genes were similar to those of known genes, 62% of which (54 genes) were categorized as proteins related to cell signaling/communication, cell structure/motility and nucleic acid management. The expression profiles in 10 human tissues of all the clones characterized in this study were examined by reverse transcription-coupled polymerase chain reaction and the chromosomal locations of the clones were determined by using human-rodent hybrid panels.

Brain Chemistry↗

Analysis of cloned cDNA and genomic sequences for phytochrome: complete amino acid sequences for two gene products expressed in etiolated Avena.

Cloned cDNA and genomic sequences have been analyzed to deduce the amino acid sequence of phytochrome from etiolated Avena. Restriction endonuclease site polymorphism between clones indicates that at least four phytochrome genes are expressed in this tissue. Sequence analysis of two complete and one partial coding region shows approximately 98% homology at both the nucleotide and amino acid levels, with the majority of amino acid changes being conservative. High sequence homology is also found in the 5'-untranslated region but significant divergence occurs in the 3'-untranslated region. The phytochrome polypeptides are 1128 amino acid residues long corresponding to a molecular mass of 125 kdaltons. The known protein sequence at the chromophore attachment site occurs only once in the polypeptide, establishing that phytochrome has a single chromophore per monomer covalently linked to Cys-321. Computer analyses of the amino acid sequences have provided predictions regarding a number of structural features of the phytochrome molecule.

Amino Acid Sequence↗

A sequence-specific single-strand DNA binding protein that contacts repressor sequences in the human GM-CSF promoter.

NF-GMb is a nuclear factor that binds to the proximal promoter of the human granulocyte-macrophage colony stimulating factor (GM-CSF) gene. NF-GMb has a subunit molecular weight of 22 kDa, is constitutively expressed in embryonic fibroblasts and binds to sequences within the adjacent CK-1 and CK-2 elements (CK-1/CK-2 region), located at approximately -100 in the GM-CSF gene promoter. These elements are conserved in haemopoietic growth factor (HGF) genes. NF-GMb binding requires the presence of repeated 5'CAGG3' sequences that overlap the binding sites for positive activators. Surprisingly, NF-GMb was found to bind solely to single-strand DNA, namely the non-coding strand of the GM-CSF CK-1/CK-2 region. NF-GMb may belong to a family of single-strand DNA binding (ssdb) proteins that have 5'CAGG3' sequences within their binding sites. Functional analysis of the proximal GM-CSF promoter revealed that sequences in the -114 to -79 region of the promoter containing the NF-GMb binding sites had no intrinsic activity in fibroblasts but could, however, repress tumour necrosis factor-alpha (TNF-alpha) inducible expression directed by downstream promoter sequences (-65 to -31). Subsequent mutation analysis showed that sequences involved in repression correlated with those required for NF-GMb binding.

Base Sequence↗

Apolipoprotein B RNA sequence 3' of the mooring sequence and cellular sources of auxiliary factors determine the location and extent of promiscuous editing.

Apolipoprotein B (apoB) RNA editing involves a cytidine to uridine transition at nucleotide 6666 (C6666) 5' of an essential cis -acting 11 nucleotide motif known as the mooring sequence. APOBEC-1 (apoB editing catalytic sub-unit 1) serves as the site-specific cytidine deaminase in the context of a multiprotein assembly, the editosome. Experimental over-expression of APOBEC-1 resulted in an increased proportion of apoB mRNAs edited at C6666, as well as editing of sites that would otherwise not be recognized (promiscuous editing). In the rat hepatoma McArdle cell line, these sites occurred predominantly 5' of the mooring sequence on either rat or human apoB mRNA expressed from transfected cDNA. In comparison, over-expression of APOBEC-1 in HepG2 (HepG2-APOBEC) human hepatoma cells, induced promiscuous editing primarily 5' of the mooring sequence, but sites 3' of the C6666 were also used more efficiently. The capacity for promiscuous editing was common to rat, rabbit and human sources of APOBEC-1. The data suggested that differences in the distribution of promiscuous editing sites and in the efficiency of their utilization may reflect cell-type-specific differences in auxiliary proteins. Deletion of the mooring sequence abolished editing at the wild type site and markedly reduced, but did not eliminate, promiscuous editing. In contrast, deletion of a pair of tandem UGAU motifs 3' of the mooring sequence in human apoB mRNA selectively reduced promiscuous editing, leaving the efficiency of editing at the wild type site essentially unaffected. ApoB RNA constructs and naturally occurring mRNAs such as NAT-1 (novel APOBEC-1 target-1) that lack this downstream element were not promiscuously edited in McArdle or HepG2 cells. These findings underscore the importance of RNA sequences and the cellular context of auxiliary factors in regulating editing site utilization.

APOBEC-1 Deaminase↗

The amino acid sequence of the 9 kDa polypeptide and partial amino acid sequence of the 20 kDa polypeptide of mitochondrial NADH:ubiquinone oxidoreductase.

Mitochondrial NADH:ubiquinone oxidoreductase (complex I) is the most complicated enzyme in the respiratory chain and is composed of at least 26 distinct polypeptides. Two hydrophilic subfractions of bovine heart complex I were systematically resolved into individual polypeptides by chromatography. Three polypeptides (51, 24, and 9 kDa) were isolated from the flavoprotein fraction (FP) of complex I, and the complete amino acid sequence of the 9 kDa polypeptide was determined. The 9 kDa polypeptide is composed of 75 amino acids with a molecular weight of 8,437. This protein exhibits no obvious sequence similarity to other proteins. The iron-sulfur protein fraction (IP) of complex I was separated into eight polypeptides, 75, 49, 30, 20, 18, 15, 13 kDa-A, and 13 kDa-B. The 20 kDa polypeptide was recognized as a novel component of IP for the first time. The N-terminal and several peptide sequences of the 20 kDa polypeptide were determined. Comparison of the sequences revealed significant sequence similarities of the 20 kDa polypeptide to the psbG gene products encoded in the chloroplast genome. The conserved sequence in these proteins was also found in the small subunit of the nickel-containing hydrogenases. These results suggest that complex I is related to other redox enzyme complexes.

Amino Acid Sequence↗

Comparative phylogenetic analyses of members of the order Planctomycetales and the division Verrucomicrobia: 23S rRNA gene sequence analysis supports the 16S rRNA gene sequence-derived phylogeny.

Almost complete 23S rRNA gene sequences were obtained from 13 planctomycete strains, the fimbriated, prosthecate bacterium Verrucomicrobium spinosum and two strains of the genus Prosthecobacter. The 23S rRNA genes were amplified by the PCR, using modified primers. The majority of the planctomycete strains investigated were shown to have 23S rRNA genes that were not linked to the 16S rRNA genes. Amplification of the 5'-termini of these genes was achieved using a novel primer-design strategy. Comparative phylogenetic analyses were performed using the 23S rRNA gene sequences determined in this study and previously determined 16S rRNA gene sequences. The phylogenetic dendrograms constructed from both datasets showed that the planctomycetes form a coherent group and distinct lineage within the domain Bacteria. Analysis of 23S rRNA gene sequences of Verrucomicrobium spinosum, Prosthecobacter fusiformis and Prosthecobacter sp. strain FC-2 showed that these organisms cluster together, as was also shown here and previously by analysis of 16S rRNA gene sequences. The distinct phylogenetic position of the division Verrucomicrobia was also supported by analysis of the 23S rRNA gene sequences, and no statistically significant phylogenetic relationship between the division Verrucomicrobia and the planctomycetes was found. The analyses presented in this study also provide further evidence that the chlamydiae are no more related to members of the order Planctomycetales and the division Verrucomicrobia than they are to members of other bacterial lineages.

Bacteria↗

Sequence analysis of the matrix protein gene of human parainfluenza virus type 3: extensive sequence homology among paramyxoviruses.

The sequences of the human parainfluenza virus type 3 (PIV3) matrix (M) mRNA [1150 nucleotides exclusive of poly(A)] and predicted M protein (353 amino acids) were determined by sequence analysis of cloned cDNA and viral genomic RNA. The gene-end sequence of the M gene differed from the semi-conserved gene-end sequence of the other PIV3 genes by an apparent insertion of eight nucleotides. The PIV3 M protein shared high sequence homology with Sendai virus and moderate homology with measles virus and canine distemper virus. Statistical analysis of the available sequences showed that the M protein was the most highly conserved parainfluenza viral protein.

Amino Acid Sequence↗

Integrated mapping, chromosomal sequencing and sequence analysis of Cryptosporidium parvum.

The apicomplexan Cryptosporidium parvum is one of the most prevalent protozoan parasites of humans. We report the physical mapping of the genome of the Iowa isolate, sequencing and analysis of chromosome 6, and approximately 0.9 Mbp of sequence sampled from the remainder of the genome. To construct a robust physical map, we devised a novel and general strategy, enabling accurate placement of clones regardless of clone artefacts. Analysis reveals a compact genome, unusually rich in membrane proteins. As in Plasmodium falciparum, the mean size of the predicted proteins is larger than that in other sequenced eukaryotes. We find several predicted proteins of interest as potential therapeutic targets, including one exhibiting similarity to the chloroquine resistance protein of Plasmodium. Coding sequence analysis argues against the conventional phylogenetic position of Cryptosporidium and supports an earlier suggestion that this genus arose from an early branching within the Apicomplexa. In agreement with this, we find no significant synteny and surprisingly little protein similarity with Plasmodium. Finally, we find two unusual and abundant repeats throughout the genome. Among sequenced genomes, one motif is abundant only in C. parvum, whereas the other is shared with (but has previously gone unnoticed in) all known genomes of the Coccidia and Haemosporida. These motifs appear to be unique in their structure, distribution and sequences.

Animals↗

Rapid characterization of HIV-1 sequence diversity using denaturing gradient gel electrophoresis and direct automated DNA sequencing of PCR products.

A direct method for visualization and isolation of sequence variants of human immunodeficiency virus type 1 (HIV-1) utilizing denaturing gradient gel electrophoresis (DGGE) combined with automated direct DNA sequencing was developed. Two fragments from the env gene and one from the nef gene of HIV-1, which together constitute approximately 1.0 kb of sequence, were amplified by PCR and analyzed. HIV-1 variants from each region were resolved and excised from the gel; this was followed by direct sequencing of different viral variants. In 9 infected patients, a limited number of dominant sequence variants could be seen in the three regions, together with a faint background of minor variants. The use of DGGE makes it possible to obtain a direct estimate of overall HIV-1 sequence diversity within patient samples without an intermediate DNA cloning step.

Base Sequence↗

Definition of the tempo of sequence diversity across an alignment and automatic identification of sequence motifs: Application to protein homologous families and superfamilies.

It is often possible to identify sequence motifs that characterize a protein family in terms of its fold and/or function from aligned protein sequences. Such motifs can be used to search for new family members. Partitioning of sequence alignments into regions of similar amino acid variability is usually done by hand. Here, I present a completely automatic method for this purpose: one that is guaranteed to produce globally optimal solutions at all levels of partition granularity. The method is used to compare the tempo of sequence diversity across reliable three-dimensional (3D) structure-based alignments of 209 protein families (HOMSTRAD) and that for 69 superfamilies (CAMPASS). (The mean alignment length for HOMSTRAD and CAMPASS are very similar.) Surprisingly, the optimal segmentation distributions for the closely related proteins and distantly related ones are found to be very similar. Also, optimal segmentation identifies an unusual protein superfamily. Finally, protein 3D structure clues from the tempo of sequence diversity across alignments are examined. The method is general, and could be applied to any area of comparative biological sequence and 3D structure analysis where the constraint of the inherent linear organization of the data imposes an ordering on the set of objects to be clustered.

Amino Acid Motifs↗

Characterization and distribution of two insertion sequences, IS1191 and iso-IS981, in Streptococcus thermophilus: does intergeneric transfer of insertion sequences occur in lactic acid bacteria co-cultures?

A chromosomal repeated sequence from Streptococcus thermophilus was identified as a new insertion sequence (IS), IS1191. This is the first IS element characterized in this species. This 1313 bp element has 28 bp imperfect terminal inverted repeats and is flanked by short direct repeats of 8 bp. The single large open reading frame of IS1191 encodes a 391-amino-acid protein which displays homologies with transposases encodes by IS1201 from Lactobacillus helveticus (44.5% amino-acid sequence identity) and by the other ISs of the IS256 family. One of the copies of IS1191 is inserted into a truncated iso-IS981 element. The nucleotide sequences of two truncated iso-IS981s from S. thermophilus and the sequence of IS981 element from Lactococcus lactis share more than 99% identity. The distribution of these insertion sequences in L. lactis and S. thermophilus strains suggests that intergeneric transfers occur during cocultures used in the manufacture of cheese.

Bacteria↗

Sequencing-based typing of HLA-A locus using mRNA and a single locus-specific PCR followed by cycle-sequencing with AmpliTaq DNA polymerase, FS.

The large number (59) of alleles now known at the HLA-A locus is a serious challenge to the existing methods for HLA typing, including many of the DNA based methods. Here, we describe a sequencing-based typing (SBT) protocol for typing of HLA-A alleles using a single A-locus-specific PCR. This reaction amplifies an 824 base pair product from cDNA, prepared from mRNA, covering exons 1-3 and most of exon 4. This product allows identification of all possible combinations of two alleles from this locus. The sequencing strategy used for allele assignment contains several improvements compared to those previously published. The enzyme AmpliTaq DNA Polymerase, FS, used, combines high-quality sequencing, i.e. long reads, low background, and uniform peak heights making the identification of heterozygous positions very reliable in a fast and easy protocol developed by determining the optima for a number of variables. Thus, this strategy meets most of the requirements for the use of sequencing in HLA typing. Furthermore, this method is very flexible. The use of a PCR primer-pair tailed with the recognition sites for two different sequencing primers allows the application of the same sets of fluorescent-labelled sequencing primers regardless of the amplified locus. Thus, the protocol can very easily be extended to cover the B- and C-locus too, simply by adding PCR reactions specific for these loci to the protocol. Using this protocol, we investigated a total of 65 cell lines and clinical samples, many of the latter chosen from samples difficult to type by serology. Our method gave in all cases unambiguous results and proved functional for work requiring the highest resolution.

Alleles↗

Nucleotide sequence of yeast gene CP A1 encoding the small subunit of arginine-pathway carbamoyl-phosphate synthetase. Homology of the deduced amino acid sequence to other glutamine amidotransferases.

A yeast DNA fragment carrying the gene CP A1 encoding the small subunit of the arginine pathway carbamoyl-phosphate synthetase has been sequenced. Only one continuous coding sequence on this fragment was long enough to account for the presumed molecular mass of CP A1 protein product. It codes for a polypeptide of 411 amino acids having a relative molecular mass, Mr, of 45 358 and showing extensive homology with the product of carA, the homologous Escherichia coli gene. CP A1 and carA products are glutamine amidotransferases which bind glutamine and transfer its amide group to the large subunits where it is used for the synthesis of carbamoyl-phosphate. A comparison of the amino acid sequences of CP A1 polypeptide with the glutamine amidotransferase domains of anthranilate and p-amino-benzoate synthetases from various sources has revealed the presence in each of these sequences of three highly conserved regions of 8, 11 and 6 amino acids respectively. The 11-residue oligopeptide contains a cysteine which is considered as the active-site residue involved in the binding of glutamine. The distances (number of amino acid residues) which separate these homology regions are accurately conserved in these various enzymes. These observations provide support for the hypothesis that these synthetases have arisen by the combination of a common ancestral glutamine amidotransferase subunit with distinct ammonia-dependent synthetases. Little homology was detected with the amide transfer domain of glutamine phosphoribosyldiphosphate amidotransferase which may be the result of a convergent evolutionary process. The flanking regions of gene CP A1 have been sequenced, 803 base pairs being determined on the 5' side and 382 on the 3' side. Several features of the 5'-upstream region of CP A1 potentially related to the control of its expression have been noticed including the presence of two copies of the consensus sequence d(T-G-A-C-T-C) previously identified in several genes subject to the general control of amino acid biosynthesis.

Amino Acid Sequence↗

BLAST 2 Sequences, a new tool for comparing protein and nucleotide sequences.

'BLAST 2 Sequences', a new BLAST-based tool for aligning two protein or nucleotide sequences, is described. While the standard BLAST program is widely used to search for homologous sequences in nucleotide and protein databases, one often needs to compare only two sequences that are already known to be homologous, coming from related species or, e.g. different isolates of the same virus. In such cases searching the entire database would be unnecessarily time-consuming. 'BLAST 2 Sequences' utilizes the BLAST algorithm for pairwise DNA-DNA or protein-protein sequence comparison. A World Wide Web version of the program can be used interactively at the NCBI WWW site (http://www.ncbi.nlm.nih.gov/gorf/bl2.++ +html). The resulting alignments are presented in both graphical and text form. The variants of the program for PC (Windows), Mac and several UNIX-based platforms can be downloaded from the NCBI FTP site (ftp://ncbi.nlm.nih.gov).

Algorithms↗

Alignment of rat cardionatrin sequences with the preprocardionatrin sequence from complementary DNA.

Mammalian atria contain peptides that promote the excretion of salt and water from the kidney. When rat atrial tissue is extracted under conditions known to inhibit proteolysis, four natriuretic peptides, cardionatrins I to IV, are consistently isolated. These peptides derive from a common precursor, preprocardionatrin, of 152 amino acids, whose sequence was determined by DNA sequencing of a complementary DNA clone. Amino acid sequencing located the start points of cardionatrins I, III, and IV in the overall sequence. Cardionatrin IV most closely resembles procardionatrin because it begins immediately after the signal sequence at residue 25. Cardionatrin III begins at residue 73, and cardionatrin I, sequenced previously, begins at residue 123. Compositional analysis indicated that each of these cardionatrins extends up to tyrosine at position 150 but lacks the terminal two arginine residues.

Amino Acid Sequence↗