Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

A linear programming approach for identifying a consensus sequence on DNA sequences.

MOTIVATION: Maximum-likelihood methods for solving the consensus sequence identification (CSI) problem on DNA sequences may only find a local optimum rather than the global optimum. Additionally, such methods do not allow logical constraints to be imposed on their models. This study develops a linear programming technique to solve CSI problems by finding an optimum consensus sequence. This method is computationally more efficient and is guaranteed to reach the global optimum. The developed method can also be extended to treat more complicated CSI problems with ambiguous conserved patterns. RESULTS: A CSI problem is first formulated as a non-linear mixed 0-1 optimization program, which is then converted into a linear mixed 0-1 program. The proposed method provides the following advantages over maximum-likelihood methods: (1) It is guaranteed to find the global optimum. (2) It can embed various logical constraints into the corresponding model. (3) It is applicable to problems with many long sequences. (4) It can find the second and the third best solutions. An extension of the proposed linear mixed 0-1 program is also designed to solve CSI problems with an unknown spacer length between conserved regions. Two examples of searching for CRP-binding sites and for FNR-binding sites in the Escherichia coli genome are used to illustrate and test the proposed method. AVAILABILITY: A software package, Global Site Seer for the Microsoft Windows operating system is available by http://www.iim.nctu.edu.tw/~cjfu/gss.htm

Algorithms↗

Construction of a contiguous 874-kb sequence of the Escherichia coli -K12 genome corresponding to 50.0-68.8 min on the linkage map and analysis of its sequence features.

The contiguous 874.423 base pair sequence corresponding to the 50.0-68.8 min region on the genetic map of the Escherichia coli K-12 (W3110) was constructed by the determination of DNA sequences in the 50.0-57.9 min region (360 kb) and two large (100 kb in all) and five short gaps in the 57.9-68.8 min region whose sequences had been registered in the DNA databases. We analyzed its sequence features and found that this region contained at least 894 potential open reading frames (ORFs), of which 346 (38.7%) were previously reported, 158 (17.7%) were homologous to other known genes, 232 (26.0%) were identical or similar to hypothetical genes registered in databases, and the remaining 158 (17.7%) showed no significant similarity to any other genes. A homology search of the ORFs also identified several new gene clusters. Those include two clusters of fimbrial genes, a gene cluster of three genes encoding homologues of the human long chain fatty acid degradation enzyme complex in the mitochondrial membrane, a cluster of at least nine genes involved in the utilization of ethanolamine, a cluster of the secondary set of 11 hyc genes participating in the formate hydrogenlyase reaction and a cluster of five genes coding for the homologues of degradation enzymes for aromatic hydrocarbons in Pseudomonas putida. We also noted a variety of novel genes, including two ORFs, which were homologous to the putative genes encoding xanthine dehydrogenase in the fly and a protein responsible for axonal guidance and outgrowth of the rat, mouse and nematode. An isoleucine tRNA gene, designated ileY, was also newly identified at 60.0 min.

Base Sequence↗

Prediction of the coding sequences of unidentified human genes. IX. The complete sequences of 100 new cDNA clones from brain which can code for large proteins in vitro.

As an extension of a series of projects for sequencing human cDNA clones derived from relatively long transcripts, we herein report the entire sequences of 100 newly determined cDNA clones with the potential of coding for large proteins in vitro. The cDNA clones were isolated from size-fractionated human brain cDNA libraries with insert sizes between 4.5 and 8.3 kb. The sequencing of these clones revealed that the average size of the cDNA inserts and of their open reading frames was 5.3 kb and 2.8 kb (930 amino acid residues), respectively. Homology search against public databases indicated that the predicted coding sequences of 86 clones exhibited significant similarities to known genes; 51 of them (59%) were related to those for cell signaling/communication, nucleic acid management, and cell structure/motility. All the clones characterized in this study are accompanied by their expression profiles in 14 human tissues examined by reverse transcription-coupled polymerase chain reaction and the chromosomal mapping data.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. XII. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

In this paper, we report the sequences of 100 cDNA clones newly determined from a set of size-fractionated human brain cDNA libraries and predict the coding sequences of the corresponding genes, named KIAA0819 to KIAA0918. These cDNA clones were selected on the basis of their coding potentials of large proteins (50 kDa and more) by using in vitro transcription/translation assays. The sequence data showed that the average sizes of the inserts and corresponding open reading frames are 4.4 kb and 2.5 kb (831 amino acid residues), respectively. Homology and motif/domain searches against the public databases indicated that the predicted coding sequences of 83 genes were similar to those of known genes, 59% of which (49 genes) were categorized as coding for proteins functionally related to cell signaling/communication, cell structure/motility and nucleic acid management. The chromosomal locations and the expression profiles of all the genes were also examined. For 54 clones including brain-specific ones, the mRNA levels were further examined among 8 brain regions (amygdala, corpus callosum, cerebellum, caudate nucleus, hippocampus, substantia nigra, subthalamic nucleus, and thalamus), spinal cord, and fetal brain.

Amino Acid Sequence↗

Determination of the cis sequence involved in catabolite repression of the Bacillus subtilis gnt operon; implication of a consensus sequence in catabolite repression in the genus Bacillus.

The mechanism underlying catabolite repression in Bacillus species remains unsolved. The gluconate (gnt) operon of Bacillus subtilis is one of the catabolic operons which is under catabolite repression. To identify the cis sequence involved in catabolite repression of the gnt operon, we performed deletion analysis of a DNA fragment carrying the gnt promoter and the gntR gene, which had been cloned into the promoter probe vector, pWP19. Deletion of the region upstream of the gnt promoter did not affect catabolite repression. Further deletion analysis of the gnt promoter and gntR coding region was carried out after restoration of promoter activity through the insertion of internal constitutive promoters of the gnt operon before the gntR gene (P2 and P3). These deletions revealed that the cis sequence involved in catabolite repression of the gnt operon is located between nucleotide positions +137 and +148. This DNA segment contains a sequence, ATTGAAAG, which may be implicated as a consensus sequence involved in catabolite repression in the genus Bacillus.

Bacillus subtilis↗

Sequence similarity of putative transposases links the maize Mutator autonomous element and a group of bacterial insertion sequences.

The Mutator transposable element system of maize is the most active transposable element system characterized in higher plants. While Mutator has been used to generate and tag thousands of new maize mutants, the mechanism and regulation of its transposition are poorly understood. The Mutator autonomous element, MuDR, encodes two proteins: MURA and MURB. We have detected an amino acid sequence motif shared by MURA and the putative transposases of a group of bacterial insertion sequences. Based on this similarity we believe that MURA is the transposase of the Mutator system. In addition we have detected two rice cDNAs in genbank with extensive similarity to MURA. This sequence similarity suggests that a Mutator-like element is present in rice. We believe that Mutator, a group of bacterial insertion sequences, and an uncharacterized rice transposon represent members of a family of transposable elements.

Amino Acid Sequence↗

V(D)J recombination frequency is affected by the sequence interposed between a pair of recombination signals: sequence comparison reveals a putative recombinational enhancer element.

The immunoglobulin heavy chain intron enhancer (Emu) not only stimulates transcription but also V(D)J recombination of chromosomally integrated recombination substrates. We aimed at reproducing this effect in recombination competent cells by transient transfection of extrachromosomal substrates. These we prepared by interposing between the recombination signal sequences (RSS) of the plasmid pBlueRec various fragments, including Emu, possibly affecting V(D)J recombination. Our work shows that sequences inserted between RSS 23 and RSS 12, with distances from their proximal ends of 26 and 284 bp respectively, can markedly affect the frequency of V(D)J recombination. We report that the entire Emu, the Emu core as well as its flanking 5' and 3' matrix associated regions (5' and 3' MARs) upregulate V(D)J recombination while the downstream section of the 3' MAR of Emu does not. Also, prokaryotic sequences markedly suppress V(D)J recombination. This confirms previous results obtained with chromosomally integrated substrates, except for the finding that the full length 3' MAR of Emu stimulates V(D)J recombination in an episomal but not in a chromosomal context. The fact that other MARs do not share this activity suggests that the effect is no mediated through attachment of the recombination substrate to a nuclear matrix-associated recombination complex but through cis-activation. The presence of a 26 bp A-T-rich sequence motif in the 5' and 3' MARs of Emu and in all of the other upregulating fragments investigated, leads us to propose that the motif represents a novel recombinational enhancer element distinct from those constituting the Emu core.

Animals↗

Nucleotide sequence and characteristics of the gene for L-lactate dehydrogenase of Thermus aquaticus YT-1 and the deduced amino acid sequence of the enzyme.

The gene for L-lactate dehydrogenase (LDH) from Thermus aquaticus YT-1 was cloned in Escherichia coli, using the Thermus caldophilus LDH gene as a hybridization probe, and its complete nucleotide sequence was determined. The LDH gene comprised 930 base pairs, starting with a GTG initiation codon. Its sequence had high homology (85.8% identity) with the LDH gene of T. caldophilus. The G + C content of the T. aquaticus gene was 70.9%, higher than that of the chromosomal DNA (67.4%). In particular, that in the third position of the codons used was 91.0%, similar to the T. caldophilus gene. The primary structure of T. aquaticus LDH was deduced from the nucleotide sequence of the LDH gene. It comprises 310 amino acid residues, as does T. caldophilus LDH, and its molecular mass was calculated to be 33,210 daltons. The amino acid sequence of the T. aquaticus LDH had 87.1% identity with that of the T. caldophilus LDH. At 23 positions, the respective residues differed in charge and polarity. These differences must be related to the differences in kinetic properties between the two enzymes. The constructed plasmid overproduced the T. aquaticus LDH in E. coli.

Amino Acid Sequence↗

Cloning and sequencing of cDNA coding for rabbit alpha-1-antiproteinase F: amino acid sequence comparison of alpha-1-antiproteinases of six mammals.

Rabbit liver cDNA coding for alpha-1-antiproteinase F has been isolated and sequenced. The protein sequence deduced from the nucleotide sequence consists of a 24 amino acid signal peptide and 389 amino acids of the mature polypeptide. Rabbit alpha-1-antiproteinase F showed 74 and 64% homology to human alpha-1-antiproteinase at the nucleotide and amino acid levels, respectively, but the N-terminal five amino acids are lacking in the rabbit protein. The sequences of alpha-1-antiproteinase F of rabbit, human, baboon, sheep, rat, and mouse show about 40% identity, and the reactive site (Met-Ser) is conserved. On the other hand, variable regions are located in the second half to the C-terminal as well as in the N-terminal region.

Amino Acid Sequence↗

The performance of several multiple-sequence alignment programs in relation to secondary-structure features for an rRNA sequence.

The performances of five global multiple-sequence alignment programs (CLUSTAL W, Divide and Conquer, Malign, PileUp, and TreeAlign) were evaluated using part of the animal mitochondrial small subunit (12S) rRNA molecule. Conserved sequence motifs derived from an alignment based on secondary structural information were used to score how well each program aligned a data set of five vertebrate and five invertebrate taxa over a range of parameter values. All of the programs could align the motifs with reasonable accuracy for at least one set of parameter conditions, although if the whole sequence was considered, similarity to the structural alignment was only 25%-34%. Use of small gap costs generally gave more accurate results, although Malign and TreeAlign generated longer alignments when gap costs were low. The programs differed in the consistency of the alignments when gap cost was varied; CLUSTAL W, Divide and Conquer, and TreeAlign were the most accurate and robust, while PileUp performed poorly as gap cost values increased, and the accuracy of Malign fluctuated. Default settings for the programs did not give the best results, and attempting to select similar parameter values in different programs did not always result in more similar alignments. Poor alignment of even well-conserved motifs can occur if these are near sites with insertions or deletions. Since there is no a priori way to determine gap costs and because such costs can vary over the gene, alignment of rRNA sequences, particularly the less well conserved regions, should be treated carefully and aided by secondary structure and conserved motifs. Some motifs are single bases and so are often invisible to alignment programs. Our tests involved the most conserved regions of the 12S rRNA gene, and alignment of less well conserved regions will be more problematical. None of the alignments we examined produced a fully resolved phylogeny for the data set, indicating that this portion of 12S rRNA is insufficient for resolution of distant evolutionary relationships.

Algorithms↗

Nucleotide sequence analysis of a matrix and small hydrophobic protein dicistronic mRNA of bovine respiratory syncytial virus demonstrates extensive sequence divergence of the small hydrophobic protein from that of human respiratory syncytial virus.

The nucleotide and deduced amino acid sequences of the matrix (M) and small hydrophobic (SH) proteins of bovine respiratory syncytial virus (BRSV) have been determined from a dicistronic mRNA. Comparison of these sequences with the corresponding published sequences of human respiratory syncytial virus (HRSV) revealed extensive overall homology at both the nucleotide and amino acid levels in the M protein, but low overall homology at both the nucleotide and amino acid levels in the SH protein. There was only 16 to 22% identity between the BRSV SH protein and the HRSV SH proteins at the C terminus. There were also an additional eight amino acids at the C terminus of BRSV. Despite the low level of identity, there were similarities in the predicted hydropathy profiles of BRSV and HRSV SH proteins. The transcription start and stop signals, which are conserved among HRSV mRNAs, were also identified in the M-SH dicistronic mRNA of BRSV. In addition, the intergenic sequence for the M-SH gene junction of BRSV was determined.

Amino Acid Sequence↗

Sequence analysis of human T cell lymphotropic virus type I strains from southern India: gene amplification and direct sequencing from whole blood blotted onto filter paper.

Human T cell lymphotropic virus type I (HTLV-I) infection in India has been found to be associated with adult T cell leukaemia/lymphoma (ATLL) and HTLV-I-associated myelopathy/tropical spastic paraparesis (HAM/TSP) among life-long residents of southern India. To examine the heterogeneity of HTLV-I strains from southern India and to determine their relationship with the sequence variants of HTLV-I from Melanesia, 1149 nucleotides spanning selected regions of the HTLV-I gag, pol, env and pX genes were amplified and directly sequenced from DNA extracted from whole blood blotted onto filter paper and from peripheral blood mononuclear cells, obtained from one patient with HAM/TSP, two with ATLL and eight asymptomatic carriers from Andhra Pradesh, Kerala and Tamil Nadu. Sequence alignments and comparisons indicated that the 11 HTLV-I strains from southern India were 99.2% to 100% identical among themselves and 98.7% to 100% identical to the Japanese prototype HTLV-I ATK. The majority of base substitutions were transitions and silent. No frameshifts, insertions, deletions or possibly disease-specific base changes were found in the regions sequenced. The observed clustering of the Indian HTLV-I strains with those from Japan, as determined by the maximum parsimony method, suggested a common source of HTLV-I infection with subsequent parallel evolution. Amplification of DNA from blood specimens collected on filter paper may be useful for the study of other blood-borne pathogens.

Adolescent↗

Molecular cloning and sequencing of the upstream region of the major Bacillus subtilis autolysin gene: a modifier protein exhibiting sequence homology to the major autolysin and the spoIID product.

The upstream region of the N-acetylmuramoyl-L-alanine amidase gene (cwlB; a major Bacillus subtilis autolysin) was cloned into Escherichia coli by chromosome walking. Sequencing of the region showed the presence of two open reading frames, one (designated as cwbA) which starts at a UUG codon and encodes a polypeptide of 705 amino acids with an M(r) of 76,725, and the other (designated as lppX), upstream of cwbA, comprising 102 amino acids and having a signal sequence characteristic of a lipoprotein. Purification of the CwbA protein and determination of its N-terminal amino acid sequence revealed that it contains a presumed signal peptide which is processed after Ala at position 25 from the N-terminal, and that the M(r) of the mature form is 75,000. The amino acid sequences of the N-terminal and C-terminal regions of CwbA were found to be highly homologous with those of the cell wall binding domain of CwlB and the spoIID gene product, respectively. CwbA stimulated the major autolysin activity approximately threefold in vitro. These data indicate that CwbA is the modifier protein of the major autolysin reported by Herbold, D. R. & Glaser, L. (1975; Journal of Biological Chemistry 250, 1676-1682). In-frame fusion between the lppX and lacZ genes demonstrated that lppX is translated in vivo and expressed during the exponential growth phase.

Amino Acid Sequence↗

Identification and sequence analysis of IS6501, an insertion sequence in Brucella spp.: relationship between genomic structure and the number of IS6501 copies.

An insertion sequence (IS) element of Brucella ovis, named IS6501, was isolated and its complete nucleotide sequence determined. IS6501 is 836 bp in length and occurs 20-35 times in the B. ovis genome and 5-15 times in other Brucella species. Analysis of the junctions at the sites of insertion revealed a small target site duplication of four bases and inverted repeats of 17 bp with one mismatch. IS6501 presents significant similarity (53.4%) with IS427 identified in Agrobacterium tumefaciens, suggesting a common ancestral sequence. A long ORF of 708 bp was identified encoding a protein with a predicted molecular mass of 26 kDa and sharing sequence identity with the hypothetical protein 1 of A. tumefaciens and with the transposase of Mycobacterium tuberculosis. IS6501 is present in all Brucella strains we have tested. Restriction fragment length polymorphism of reference and field strains of two species (B. melitensis and B. ovis) was studied using either pulsed field gel electrophoresis (PFGE) on XbaI-digested DNA or hybridization of EcoRI-digested DNA using IS6501 as a probe. The genome of B. melitensis biovar 3 contains about 10 IS copies per genome and field strains of the same species could not be distinguished either by IS hybridization or by XbaI (PFGE) restriction patterns. In contrast, the number of IS copies in the B. ovis genome is around 30 and the different field strains can be differentiated by both methods.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Comparing low coverage random shotgun sequence data from Brassica oleracea and Oryza sativa genome sequence for their ability to add to the annotation of Arabidopsis thaliana.

Since the completion of the Arabidopsis thaliana genome sequence, there is an ongoing effort to annotate the genome as accurately as possible. Comparing genome sequences of related species complements the current annotation strategies by identifying genes and improving gene structure. A total of 595,321 Brassica oleracea shotgun reads were sequenced by TIGR (The Institute for Genome Research) and the collaboration of Washington University and Cold Spring Harbor. Vicogenta (a genome viewer based on GMOD and GBrowse) was created to view the current annotation and sequence alignments for Arabidopsis. Brassica reads were compared with the Arabidopsis genome and proteome databases using BLAST. Hypothetical genes and conserved unannotated regions on the short arm of chromosome 4 from Arabidopsis were experimentally verified using RT-PCR. We were able to improve the Arabidopsis annotation by identifying 25 genes that were missed, and confirming expression of 43 hypothetical genes in Arabidopsis. We were also able to detect conservation in genes whose transcription is normally suppressed due to methylation. We also examined how useful the O. sativa genome and ESTs from other species are, compared with Brassica, in improving the Arabidopsis annotation.

Amino Acid Sequence↗

Polymorphism of HLA-DRw52-associated DRB1 genes as defined by sequence-specific oligonucleotide probe hybridization and sequencing.

We have used group-specific DNA amplification and sequence-specific oligonucleotide probe (SSOP) hybridization to study DRB1 sequence polymorphisms associated with DR3, DRw11(5), DRw12(5), DRw13(w6), DRw14(w6) and DRw8 alleles. Group-specific amplification of DRw52-associated DRB1 alleles was achieved using a 5' amplification primer designed to hybridize with a first hypervariable region (HVR) sequence common to all known alleles in this group, together with a 3' intron primer. Prospective SSOP typing of DR3, DRw11, DRw12, DRw13, DRw14 and DRw8 alleles was performed in 318 individuals, including 124 patients, 46 family members and 148 unrelated marrow donors. Among the 395 DRw52-associated DRB1 alleles tested in our study, a subtype corresponding to the previously defined alleles DRB1*0301-2 (DR3), DRB1*1101-4 (DR5), DRB1*1201-2 (DR5), DRB1*1301-5 (DRw6), DRB1*1401-2 and 1404 (DRw6), and DRB1*0801-4 (DRw8) could be assigned in all but 6 individuals (1.9%) tested. In addition to the 22 known alleles, we identified two new DRw6-associated alleles, DRB1*13.MW(1) and DRB1*14.GB(1). DRB1*13.MW typed serologically as DRw13 and was identical to DRB1*1301 except at codon 71 where AGG encodes arginine instead of GAG encoding glutamic acid. DRB1*14.GB represents a DRB1*1402 variant whose sequence at codon 86 encodes valine (GTG) instead of glycine (GGT). These results demonstrate that SSOP methods represent an efficient and precise approach for typing DRB1 alleles and for identifying potential novel variants previously unrecognized by conventional typing methods.

Alleles↗

The aconitase of Escherichia coli. Nucleotide sequence of the aconitase gene and amino acid sequence similarity with mitochondrial aconitases, the iron-responsive-element-binding protein and isopropylmalate isomerases.

The nucleotide sequence of the aconitase gene (acn) of Escherichia coli was determined and used to deduce the primary structure of the enzyme. The coding region comprises 2670 bp (890 codons excluding the start and stop codons) which define a product having a relative molecular mass of 97,513 and an N-terminal amino acid sequence consistent with those determined previously for the purified enzyme. The acn gene is flanked by the cysB gene and a putative riboflavin biosynthesis gene resembling the ribA gene of Bacillus subtilis. The 1004-bp cysB--acn intergenic region contains several potential promoter and regulatory sequences. The amino acid sequence of the E. coli aconitase is similar to the mitochondrial aconitases (27-29% identity) and the isopropylmalate isomerases (20-21% identity) but it is most similar to the human iron-responsive-element-binding protein (53% identity). The three cysteine residues involved in ligand binding to the [4Fe-4S] centre are conserved in all of these proteins. Of the remaining 17 active-site residues assigned for porcine aconitase, 16 are conserved in both the bacterial aconitase and the iron-responsive-element-binding protein and 14 in the isopropylmalate isomerases. It is concluded that the bacterial and mitochondrial aconitases, the isopropylmalate isomerases and the iron-responsive-element-binding protein form a family of structurally related proteins, which does not include the Fe-S-containing fumarases. These relationships raise the possibility that the iron-responsive-element-binding protein may be a cytoplasmic aconitase and that the E. coli aconitase may have an iron-responsive regulatory function.

Aconitate Hydratase↗

Insertion in barnase of a loop sequence from ribonuclease T1. Investigating sequence and structure alignments by protein engineering.

Barnase was mutated by inserting into its active site loop sequences found in the related enzyme ribonuclease T1 (RNase T1), according to either structural or sequential similarity alignments. The barnase/RNase T1 hybrid corresponding to the structural alignment of the two proteins, endo-[RNaseT1-(93-99)]102abarnase, contains RNase T1 residues at positions 93-99 inserted between residues at positions 102 and 103 of barnase. The other constructed mutant, endo-[RNaseT1-(95-98)]104abarnase, has RNase T1 residues at positions 95-98 inserted between residues at positions 104 and 105 in barnase, corresponding to published sequence alignments of the two proteins in this region. The mutants were characterized by absorbance, fluorescence and CD spectroscopy; the stability, folding and unfolding kinetics, and catalytic activity were measured and compared with the wild-type enzyme. Endo-[RNaseT1-(93-99)]102abarnase, the mutant protein corresponding to the structural alignment of barnase with ribonuclease T1, shows a slightly higher stability (approximately 5 kJ/mol) towards urea and heat denaturation than the mutant endo-[RNaseT1-(95-98)]104abarnase, designed according to a sequence alignment between the two enzymes. Both mutants have very low catalytic activity, although the effect of mutation is almost entirely limited to kcat in the case of the mutant corresponding to the structural alignment between barnase and ribonuclease T1, while both kcat and Km are affected in the mutant corresponding to the sequence alignment between the two enzymes. Thus, the superiority of structural over sequential alignments cannot be supported conclusively by direct experiment in the present case.

Amino Acid Sequence↗