Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Polar zipper sequence in the high-affinity hemoglobin of Ascaris suum: amino acid sequence and structural interpretation.

The extracellular hemoglobin of Ascaris has an extremely high oxygen affinity (P50 = 0.004 mmHg). It consists of eight identical subunits of molecular weight 40,600. Their sequence, determined by protein chemistry, shows two tandemly linked globin-like sequences and an 18-residue C-terminal extension. Two N-linked glycosylation sites contain equal ratios of mannose/glucosamine/fucose of 3:2:1. Electron micrographs suggest that the eight subunits form a polyhedron of point symmetry D4, or 42. The C-terminal extension contains a repeat of the sequence Glu-Glu-His-Lys, which would form a pattern of alternate glutamate and histidine side chains on one side and of glutamate and lysine side chains on the other side of a beta strand. We propose that this represents a polar zipper sequence and that the C-terminal extensions are joined in an eight-stranded beta barrel at the center of the molecule, with histidine and glutamate side chains inside and lysine and glutamate side chains outside the barrel compensating each other's charges. The amino acid sequence of Ascaris hemoglobin fails to explain its high oxygen affinity.

Amino Acid Sequence↗

Adipose pyruvate carboxylase: amino acid sequence and domain structure deduced from cDNA sequencing.

The complete amino acid sequence of 3T3-L1 adipocyte pyruvate carboxylase (PC) [pyruvate:carbon-dioxide ligase (ADP-forming), EC 6.4.1.1] has been deduced from sequencing overlapping cDNA clones obtained from an adipocyte cDNA library constructed in the lambda Zap vector. The encoding mRNA for PC promoter contains 4067 nt, including a 3534-nt coding sequence and noncoding regions of 100 and 433 nt at the 5' and 3' ends, respectively. The biotinylated lysine of the encoded PC promoter (1178 amino acids with a calculated M(r) of apocarboxylase = 129,784) is located 35 residues from the COOH-terminal end and, as in most other biotin enzymes, is in the consensus sequence AMKM. The adipocyte PC is closely similar (53% identity) to the yeast enzyme and contains different segments that are homologous with regions from the biotin carboxylase component of Escherichia coli acetyl-CoA carboxylase, the keto acid-binding subunits of Propionibacterium shermanii oxaloacetate transcarboxylase and Klebsiella pneumoniae oxaloacetate decarboxylase, and to the biotin carboxyl-carrier protein of the bacterial biotin enzymes. In addition to the putative mitochondrial targeting signal, functional domains are readily identifiable in the sequence and are in the following order: biotin carboxylase-carboxyltransferase-biotin carboxyl-carrier protein, as proposed for yeast PC.

3T3 Cells↗

Functional analysis of the chimpanzee and human apo(a) promoter sequences: identification of sequence variations responsible for elevated transcriptional activity in chimpanzee.

Lp(a) concentrations vary considerably among individuals and are primarily determined by the apo(a) gene locus. We have previously shown that mean plasma Lp(a) levels in the chimpanzee are significantly higher than those observed in humans (Doucet, C., Huby, T., Chapman, J., and Thillet, J. (1994) J. Lipid Res 35, 263-270). To evaluate the possibility that this difference may result from a high level of expression of chimpanzee apo(a), we cloned and sequenced 1.4 kilobase (kb) of the 5'-flanking region of the gene and compared promoter activity to that of its human counterpart. Sequence analysis revealed 98% homology between chimpanzee and human apo(a) 5' sequences; among the differences observed, two involved polymorphic sites associated with Lp(a) levels in humans. The TTTTA repeat located 1.3 kb 5' of the apo(a) gene, present in a variable number of copies (n = 5-12) in humans, is uniquely present as four copies in the chimpanzee sequence. The second position concerns the +93 C>T polymorphism that creates an additional ATG start codon in the human apo(a) gene, thereby impairing translation efficiency. In chimpanzee, this position did not appear polymorphic, and a base difference at position +94 precluded the presence of an additional ATG. In transient transfection assays, the chimpanzee apo(a) promoter exhibited a 5-fold elevation in transcriptional activity as compared with its human counterpart. This marked difference in activity was maintained with either 1.4 kb of 5' sequence or the minimal promoter region -98 to +141 of the human and chimpanzee apo(a) genes. Using point mutational analyses, nucleotides present at positions -3, -2, and +8 (relative to the start site of transcription) were found to be essential for the high transcription efficiency of the chimpanzee apo(a) promoter. High transcriptional activity of the chimpanzee apo(a) gene may therefore represent a key factor in the elevated plasma Lp(a) levels characteristic of this non-human primate.

Animals↗

DivergentSet, a tool for picking non-redundant sequences from large sequence collections.

DivergentSet addresses the important but so far neglected bioinformatics task of choosing a representative set of sequences from a larger collection. We found that using a phylogenetic tree to guide the construction of divergent sets of sequences can be up to 2 orders of magnitude faster than the naive method of using a full distance matrix. By providing a user-friendly interface (available online) that integrates the tasks of finding additional sequences, building and refining the divergent set, producing random divergent sets from the same sequences, and exporting identifiers, this software facilitates a wide range of bioinformatics analyses including finding significant motifs and covariations. As an example application of DivergentSet, we demonstrate that the motifs identified by the motif-finding package MEME (Motif Elicitation by Maximum Entropy) are highly unstable with respect to the specific choice of sequences. This instability suggests that the types of sensitivity analysis enabled by DivergentSet may be widely useful for identifying the motifs of biological significance.

Amino Acid Sequence↗

Prediction whether a human cDNA sequence contains initiation codon by combining statistical information and similarity with protein sequences.

MOTIVATION: In the previous works, we developed ATGpr, a computer program for predicting the fullness of a cDNA, i.e. whether it contains an initiation codon or not. Statistical information of short nucleotide fragments was fully exploited in the prediction algorithm. However, sequence similarities to known proteins, which are becoming increasingly available due to recent rapid growth of protein database, were not used in the prediction. In this work, we present a new prediction algorithm based on both statistical and similarity information, which provides better performance in sensitivity and specificity. RESULTS: We evaluated the accuracy of ATGpr for predicting fullness of cDNA sequences from human clustered ESTs of UniGene, and we obtained specificity, sensitivity, and correlation coefficient of this prediction. Specificity and sensitivity crossed at 46% over the ATGpr score threshold of 0.33 and the maximum correlation coefficient of 0.34 was obtained at this threshold. Without ATGpr we found it effective to use alignments with known proteins for predicting the fullness of cDNA sequences. That is, specificity increased monotonously as similarity (identity of the alignments) increased. Specificity was achieved greater than 80% if identity was greater than 40%. For more effective prediction of fullness of cDNA sequences we combined the similarity (identity of query sequence) with known proteins and ATGpr score. As a result, specificity became greater than 80% if identity was greater than 20%. AVAILABILITY: The prediction program, called ATGpr_ sim, is available at http://www.hri.co.jp/atgpr/ATGpr_sim.html CONTACT: nisikawa@crl.hitachi.co.jp

Amino Acid Sequence↗

'Size leap' algorithm: an efficient extraction of the longest common motifs from a molecular sequence set. Application to the DNA sequence reconstruction.

We propose a new method, called 'size leap' algorithm, of search for motifs of maximum size and common to two fragments at least. It allows the creation of a reduced database of motifs from a set of sequences whose size obeys the series of Fibonacci numbers. The convenience lies in the efficiency of the motif extraction. It can be applied in the establishment of overlap regions for DNA sequence reconstruction and multiple alignment of biological sequences. The method of complete DNA sequence reconstruction by extraction of the longest motifs ('anchor motifs') is presented as an application of the size leap algorithm. The details of a reconstruction from three sequenced fragments are given as an example.

Algorithms↗

Prediction of the coding sequences of unidentified human genes. XIV. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

To extend our cDNA project for accumulating basic information on unidentified human genes, we newly determined the sequences of 100 cDNA clones from a set of size-fractionated human adult and fetal brain cDNA libraries, and predicted the coding sequences of the corresponding genes, named KIAA1019 to KIAA1118. The sequencing of these clones revealed that the average size of the inserts and corresponding open reading frames were 5.0 kb and 2.6 kb (880 amino acid residues), respectively. Database search of the predicted amino acid sequences classified 58 predicted gene products into the five functional categories, such as cell signaling/communication, cell structure/motility, nucleic acid management, protein management and cell division. It was also found that, for 34 gene products, homologues were detected in the databases, which were similar in sequence through almost the entire regions. The chromosomal locations of the genes were determined by using human-rodent hybrid panels unless their mapping data were already available in the public databases. The expression profiles of all the genes among 10 human tissues, 8 brain regions (amygdala, corpus callosum, cerebellum, caudate nucleus, hippocampus, substania nigra, subthalamic nucleus, and thalamus), spinal cord, fetal brain and fetal liver were also examined by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

Repetitive sequences in the crocodilian mitochondrial control region: poly-A sequences and heteroplasmic tandem repeats.

Heteroplasmic tandem repeats in the mitochondrial control region have been documented in a wide variety of vertebrate species. We have examined the control region from 11 species in the family Crocodylidae and identified two different types of heteroplasmic repetitive sequences in the conserved sequence block (CSB) domain-an extensive poly-A tract that appears to be involved in the formation of secondary structure and a series of tandem repeats located downstream ranging from approximately 50 to approximately 80 bp in length. We describe this portion of the crocodylian control region in detail and focus on members of the family Crocodylidae. We then address the origins of the tandemly repeated sequences in this family and suggest hypotheses to explain possible mechanisms of expansion/contraction of the sequences. We have also examined control region sequences from Alligator and Caiman and offer hypotheses for the origin of tandem repeats found in those taxa. Finally, we present a brief analysis of intraindividual and interindividual haplotype variation by examining representatives of Morelet's crocodile (Crocodylus moreletii).

Alligators and Crocodiles↗

Determination of the complete nucleotide sequence of the Sendai virus genome RNA and the predicted amino acid sequences of the F, HN and L proteins.

We previously determined the 3' proximal 5,824 nucleotides of the Sendai virus genome RNA (Nucleic Acids Res. 11, 7317-7330, 1983; Nucleic Acids Res. 12, 7965-7973, 1984), and present here the sequence of the remaining 5' proximal 9,559 nucleotides. Thus, this is the first paramyxovirus to have its genome organization elucidated. The set of complementary DNA clones used was prepared by the method of Okayama and Berg from polyadenylylated viral genome RNA. We sequenced the region containing the 5' proximal half of the F gene, and the subsequent HN and L genes, and predicted the complete amino acid sequence of the products of these genes. Sequence analyses confirmed that all the genes are flanked by consensus sequences and suggest that the viral mRNAs are capable of forming stem-and-loop structures. Comparison of the F and HN glycoproteins of Sendai virus with those of simian virus 5 strongly suggests that the cysteine residues are highly important for maintenance of the molecular structures of these glycoproteins.

Amino Acid Sequence↗

Rabbit muscle creatine kinase: genomic cloning, sequencing, and analysis of upstream sequences important for expression in myocytes.

Muscle creatine kinase (MCK) is a major enzyme of cellular energy metabolism that is expressed upon differentiation of myoblasts into myotubes. Previously we cloned and sequenced the entire rabbit enzyme cDNA which was used as a probe in these studies to obtain a genomic clone from a rabbit library. The transcription start site was identified by primer extension analysis and over 800 bp of 5' flanking DNA was sequenced. Comparison of this sequence with the published sequences from the upstream regions of the mouse MCK gene and the human MCK gene showed two conserved regions and a large intervening block of non-conserved sequence. The conserved regions are separated by about 800 bp in the mouse and by about 400 bp in the human, but are much closer (200 bp) in the rabbit. The upstream conserved region of the mouse gene encompasses a region possessing the properties of an enhancer and containing two MyoD binding sites; the downstream element is adjacent to the start of transcription. A set of of overlapping deletions of the 5' upstream DNA was fused to the CAT gene and transfected into mouse C2 myocytes, chick primary myocytes, and chick primary liver cells. Constructs which contained both conserved 5' regions were strongly expressed in C2 and chick myocytes, but were not expressed (above background) in primary liver cells. Surprisingly, while the upstream enhancer element was required for strong expression in C2 myocytes, it was less important for expression in chick myocytes. This suggests that there are important muscle-specific transcriptional signals in the proximal promoter region of mammalian MCK genes.

Animals↗

CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.

The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved for the alignment of divergent protein sequences. Firstly, individual weights are assigned to each sequence in a partial alignment in order to down-weight near-duplicate sequences and up-weight the most divergent ones. Secondly, amino acid substitution matrices are varied at different alignment stages according to the divergence of the sequences to be aligned. Thirdly, residue-specific gap penalties and locally reduced gap penalties in hydrophilic regions encourage new gaps in potential loop regions rather than regular secondary structure. Fourthly, positions in early alignments where gaps have been opened receive locally reduced gap penalties to encourage the opening up of new gaps at these positions. These modifications are incorporated into a new program, CLUSTAL W which is freely available.

Algorithms↗

The sequence of the stem and flanking sequences at the 3' end of histone mRNA are critical determinants for the binding of the stem-loop binding protein.

Complexes of different electrophoretic mobility containing the stem-loop binding protein, a 45 kDa protein, bound to the stem-loop at the 3' end of histone mRNA, are present in both nuclear and cytoplasmic extracts from mammalian cells. We have determined the effect of changes in the loop, in the stem and in the flanking sequences on the affinity of the SLBP for the 3' end of histone mRNA. The sequence of the stem is particularly critical for SLBP binding. Specific sequences both 5' and 3' of the stem-loop are also required for high-affinity binding. Expanding the four base loop by one or two uridines reduced but did not abolish SLBP binding. RNA footprinting experiments show that the flanking sequences on both sides of the stem-loop are critical for efficient binding, but that cleavages in the loop do not abolish binding. Thus all three regions of the RNA sequence contribute to SLBP binding, suggesting that the 26 nt at the 3' end of histone mRNA forms a defined tertiary structure recognized by the SLBP.

Animals↗

Nucleotide sequence of a 28-kbp portion of rice mitochondrial DNA: the existence of many sequences that correspond to parts of mitochondrial genes in intergenic regions.

The nucleotide sequence of a 27,588-bp region of rice mitochondrial DNA was determined. This sequence contains putative genes that encode initiator methionine tRNA (trnfM), subunits III (nad3) and IV (nad4) of the NADH dehydrogenase complex, and ribosomal proteins S3 (rps3), S12 (rps12) and L16 (rpl16). An open reading frame that contains sequences homologous to parts of rps2 and atpA is also present. In addition to these regions, there are many short sequences with homology to fragments of mitochondrial DNAs from rice or other plants. These sequences may be remnants of multiple rearrangements of the genome and their presence seems to explain, in part, the large sizes of the mitochondrial genomes of higher plants.

Base Sequence↗

Analysis of the constitution of the beer yeast genome by PCR, sequencing and subtelomeric sequence hybridization.

The lager brewing yeasts, Saccharomyces pastorianus (synonym Saccharomyces carlsbergensis), are allopolyploid, containing parts of two divergent genomes. Saccharomyces cerevisiae contributed to the formation of these hybrids, although the identity of the other species is still unclear. The presence of alleles specific to S. cerevisiae and S. pastorianus was tested for by PCR/RFLP in brewing yeasts of various origins and in members of the Saccharomyces sensu stricto complex. S. cerevisiae-type alleles of two genes, HIS4 and YCL008c, were identified in another brewing yeast, S. pastorianus CBS 1503 (Saccharomyces monacensis), thought to be the source of the other contributor to the lager hybrid. This is consistent with the hybridization of S. cerevisiae subtelomeric sequences X and Y' to the electrophoretic karyotype of this strain. S. pastorianus CBS 1503 (S. monacensis) is therefore probably not an ancestor of S. pastorianus, but a related hybrid. Saccharomyces bayanus, also thought to be one of the contributors to the lager yeast hybrid, is a heterogeneous taxon containing at least two subgroups, one close to the type strain, CBS 380T, the other close to CBS 395 (Saccharomyces uvarum). The partial sequences of several genes (HIS4, MET10, URA3) were shown to be identical or very similar (over 99%) in S. pastorianus CBS 1513 (S. carlsbergensis), S. bayanus CBS 380T and its close derivatives, showing that S. pastorianus and S. bayanus have a common ancestor. A distinction between two subgroups within S. bayanus was made on the basis of sequence analysis: the subgroup represented by S. bayanus CBS 395 (S. uvarum) has 6-8% sequence divergence within the genes HIS4, MET10 and MET2 from S. bayanus CBS 380T, indicating that the two S. bayanus subgroups diverged recently. The detection of specific alleles by PCR/RFLP and hybridization with S. cerevisiae subtelomeric sequences X and Y' to electrophoretic karyotypes of brewing yeasts and related species confirmed our findings and revealed substantial heterogeneity in the genome constitution of Czech brewing yeasts used in production.

Alleles↗

Partial nucleotide sequence and deduced amino acid sequence of the structural proteins of dengue virus type 2, New Guinea C and PUO-218 strains.

The nucleotide sequence and the deduced amino acid sequence for the genes encoding the structural proteins of two strains of dengue virus type 2 (DEN-2) were determined from cDNA clones. The genes for C, prM(M) and E proteins were sequenced for the prototype DEN-2 virus, the New Guinea C strain. Also sequenced were the prM(M) and E genes of PUO-218. This strain of DEN-2 was isolated during 1980 in Bangkok and had received a limited number of laboratory passages. Comparisons of the newly determined sequences with those published for the Jamaica 1409 and Puerto Rico PR-159 (S1 vaccine candidate) strains revealed a close relationship between New Guinea C virus and both the Jamaica and PUO-218 viruses (greater than 96% similarity in nucleotides of the E gene), whereas S1 virus was the most divergent.

Amino Acid Sequence↗

Intratype sequence variation among clinical isolates of the human papillomavirus type 6 L1 ORF: clustering of mutations and identification of a frequent amino acid sequence variant.

Human papillomavirus type 6 (HPV-6) is the causative agent of condyloma acuminata, a common sexually transmitted disease. Virus-like particles (VLPs) assembled from the L1 major capsid protein represent promising candidates for prophylactic vaccines. However, any intratype sequence variation among HPV-6 L1 ORFs will influence which sequence is used for a vaccine according to its prevalence in the population and its propensity for VLP production. Therefore, we have analysed the entire L1 nucleotide sequence of 17 clinical isolates of HPV-6 from the London area. We found 28 positions where changes from the prototype HPV-6b L1 occurred, showing that HPV-6 L1 intratype variation is greater than previously reported. The most frequently observed substitutions are clustered into three discrete regions: R1 (nt 5920-6075), R2 (nt 6590-6670) and R3 (nt 7070-7230). Indeed, most of the nucleotide substitutions within the HPV-6 L1 reported worldwide also map to these regions. The R3 region contains predominantly non-silent substitutions, the most common of which is a G-to-C substitution at position 7079. This results in a Glu-to-Gln change at aa 431, although this change had no effect on VLP yield or stability. This substitution defines a new HPV-6 L1 amino acid sequence that is more abundant in the isolates examined than any other reported sequence.

Amino Acid Sequence↗

A computer program for aligning a cDNA sequence with a genomic DNA sequence.

We address the problem of efficiently aligning a transcribed and spliced DNA sequence with a genomic sequence containing that gene, allowing for introns in the genomic sequence and a relatively small number of sequencing errors. A freely available computer program, described herein, solves the problem for a 100-kb genomic sequence in a few seconds on a workstation.

Algorithms↗

Generation of expressed sequence tags of random root cDNA clones of Brassica napus by single-run partial sequencing.

Two hundred thirty-seven expressed sequence tags (ESTs) of Brassica napus were generated by single-run partial sequencing of 197 random root cDNA clones. A computer search of these root ESTs revealed that 21 ESTs show significant similarity to the protein-coding sequences in the existing data bases, including five stress- or defense-related genes and four clones related to the genes from other kingdoms. Northern blot analysis of the 10 data base-matched cDNA clones revealed that many of the clones are expressed most abundantly in root but less abundantly in other organs. However, two clones were highly root specific. The results show that generation of the root ESTs by partial sequencing of random cDNA clones along with the expression analysis is an efficient approach to isolate genes that are functional in plant root in a large scale. We also discuss the results of the examination of cDNA libraries and sequencing methods suitable for this approach.

Amino Acid Sequence↗