Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Modulation of RGD sequence motifs regulates disintegrin recognition of alphaIIb beta3 and alpha5 beta1 integrin complexes. Replacement of elegantin alanine-50 with proline, N-terminal to the RGD sequence, diminishes recognition of the alpha5 beta1 complex with restoration induced by Mn2+ cation.

Several recent studies have demonstrated that the amino acid residues flanking the RGD sequence of high-affinity ligands modulate their specificity of interaction with integrin complexes. The present study has addressed the role of the residues flanking the RGD sequence in regulating the recognition by disintegrin of the alphaIIb beta3 and alpha5beta1 complexes by construction of a panel of recombinant molecules of Elegantin (the platelet aggregation inhibitor from the venom of Trimerasurus elegans) expressing specific RGD sequence motifs. Wild-type Elegantin (ARGDNP) and several variants including Eleg. AM (ARGDMP), Eleg. PM (PRGDMP) and Eleg. PN (PRGDNP) were expressed as glutathione S-transferase (GST) fusion proteins in Escherichia coli. The inhibitory efficacies of the panel of Elegantin variants were analysed in platelet adhesion assays with substrates immobilized with fibrinogen and fibronectin. Elegantin molecules containing an Ala residue N-terminal to the RGD sequence (wild-type Elegantin and Eleg. AM) showed strong inhibitory activity towards alphaIIbbeta3-dependent platelet adhesion on fibronectin, whereas a Pro residue in this position (Eleg. PM and Kistrin, the inhibitor from the venom of Calloselasma rhodostoma) engendered lower activity. The decreased activity could not be attributed to a decrease in the affinity of the disintegrin for the alphaIIb beta3 complex because both Eleg. AM and Eleg. PM had similar Kd (app) values. In contrast, Elegantin molecules into which a Met residue was introduced in place of the Asn residue C-terminal to the RGD sequence showed 10-13-fold elevated inhibitory activity towards platelet adhesion on fibrinogen and this was maintained with either a Pro or Ala residue N-terminal to the RGD sequence. In experiments with the alpha5 beta1 complex on K562 cells, the inhibitory efficacies of the panel of Elegantin molecules were analysed under two different cation conditions. First, in the presence of Ca2+/Mg2+, K562 cell adhesion on fibronectin was inhibited equally well by Elegantin and Eleg. AM but inhibited poorly by Eleg. PM and Kistrin. In contrast with platelets, the decreased inhibitory efficacy of the PRGDMP disintegrins was due to poor recognition of the alpha5 beta1 complex. In the presence of Mn2+ cation, K562 cell adhesion on fibrinogen was observed in an alpha5 beta1-dependent manner. Under these conditions both PRGD and ARGD containing disintegrins were strong inhibitors of K562 cell adhesion on fibrinogen and this was due to a markedly improved recognition of the alpha5 beta1 complex by the PRGD molecules. These observations demonstrate the pivotal role of the amino acids flanking the RGD sequence for disintegrin recognition of integrin complexes and highlight the subtle nature by which integrin-ligand binding specificity can be modulated by both cation and adhesive motif.

Amino Acid Sequence↗

Poisson process approximation for sequence repeats, and sequencing by hybridization.

Sequencing by hybridization is a tool to determine a DNA sequence from the unordered list of all l-tuples contained in this sequence; typical numbers for l are l = 8, 10, 12. For theoretical purposes we assume that the multiset of all l-tuples is known. This multiset determines the DNA sequence uniquely if none of the so-called Ukkonen transformations are possible. These transformations require repeats of (l-1)-tuples in the sequence, with these repeats occurring in certain spatial patterns. We model DNA as an i.i.d. sequence. We first prove Poisson process approximations for the process of indicators of all leftmost long repeats allowing self-overlap and for the process of indicators of all left-most long repeats without self-overlap. Using the Chen-Stein method, we get bounds on the error of these approximations. As a corollary, we approximate the distribution of longest repeats. In the second step we analyze the spatial patterns of the repeats. Finally we combine these two steps to prove an approximation for the probability that a random sequence is uniquely recoverable from its list of l-tuples. For all our results we give some numerical examples including error bounds.

Algorithms↗

Aligning a DNA sequence with a protein sequence.

We develop several algorithms for the problem of aligning DNA sequence with a protein sequence. Our methods account for frameshift errors, but not for introns in the DNA sequence. Thus, they are particularly appropriate for comparing a cDNA sequence that suffers from sequencing errors with an amino acid sequence or a protein sequence database. We describe algorithms for computing optimal alignments for several definitions of DNA-protein alignment, verify sufficient conditions for equivalence of certain definitions, describe techniques for efficient implementation, and discuss experience with these ideas in a new release of the FASTA suite of database-searching programs.

Algorithms↗

A procedure to verify an amino acid sequence which has been derived from a nucleotide sequence: application to the 26S RNA of Semliki Forest virus.

We describe a peptide sequencing procedure which can be used to verify an amino acid sequence which is derived from a nucleotide sequence. One first labels the protein with a 3H- and a 14C-labelled amino acid and then cleaves the protein into a set of peptides using a cleavage reaction specific for a particular amino acid residue. Finally one performs Edman degradations on the whole mixture of peptides. The released amino acids reflect the combined aminoterminal amino acid sequences of all the peptides that have been formed by the cleavage reaction. The data can therefore be used to check a deduced sequence simultaneously at several regions of the polypeptide chain. We have applied this sequencing procedure to verify the amino acid sequence deduced from the 26S RNA of Semliki Forest virus.

Amino Acid Sequence↗

Mung bean nuclease cleavage of a dA + dT-rich sequence or an inverted repeat sequence in supercoiled PM2 DNA depends on ionic environment.

We have determined the nucleotide sequences around two alternative sites cleaved in supercoiled PM2 DNA by single-strand-specific mung bean nuclease in different ionic environments. In 10 mM Tris-HC1 (pH 7.0, 37 degrees C), the major site is a dA+dT-rich sequence which maps with a known early denaturation region at 0.75 map units. About 30 cleavages occurred in a 135 bp region. Cleavages were largely excluded at (dA)n . (dT)n (n = 3-7) sequences. Cleavage patterns of this type have not been previously observed in dA+dT-rich sequences. With the addition of 0.1 M NaC1 the major alternative site occurred in a hyphenated inverted repeat sequence 500 bp away (0.70 map units) and did not map to an early denaturation region. One major and 4 minor cleavages occurred in the region between the repeats, suggesting that a hairpin containing at most a 12 bp stem and 10 base loop is recognized. The basis for nuclease recognition of the dA+dT-rich sequence is not clear. The differences in the sequences and cleavage patterns at the alternative sites indicate that their secondary structures differ.

Bacteriophages↗

Analysis of repetitive sequence elements containing tRNA-like sequences.

Several repetitive sequence elements from diverse species share extensive sequence homology with tRNA molecules. Analysis of the tRNA-like sequences within these elements suggest that they have originated from authentic tRNA sequences. Elements containing tRNA-like sequences can be divided into three distinct groups whose members share extensive sequence homology, have similar sequence organization and have unique species distribution. We suggest that these three groups represent independent examples of retroposon families that have originated from tRNAs.

Animals↗

Comparison of the sequence specificity of bleomycin cleavage in two slightly different DNA sequences.

The sequence specificity of bleomycin damage was investigated utilising 340 bp alpha-DNA (a middle repetitive sequence in the human genome) as a target sequence. The following significant facts were found:- i) The dinucleotides GT and GC were cleaved on all occasions, GA most of the time, and AT, AC, GG and AA cleaved some of the time; ii) The base immediately 5' to the purine-pyrimidine dinucleotides was found to be statistically highly significant in determining the degree of damage caused by bleomycin, while other nearest neighbour bases had no significant effect; iii) The sequence specificity of bleomycin damage was determined on both strands and it was found that damage on either strand follows the above dinucleotide preference and is independent of the extent of damage on the opposite strand; iv) Bleomycin damage was compared between genomic 340 bp alpha-DNA and a cloned alpha-DNA with eleven base substitutions relative to the "consensus" sequence. There were forty-nine detectable differences in intensity of damage between these two DNA molecules. Although four of the differences can be directly attributed to changes in base sequence, the remaining differences were not at the base substitution sites. Some of the differences were over fifty base pairs from the nearest base substitution. We propose that the majority of these differences are due to microvariation in the structure of DNA with a slightly different DNA sequence.

Base Sequence↗

Respiratory syncytial virus fusion glycoprotein: nucleotide sequence of mRNA, identification of cleavage activation site and amino acid sequence of N-terminus of F1 subunit.

The amino acid sequence of respiratory syncytial virus fusion protein (Fo) was deduced from the sequence of a partial cDNA clone of mRNA and from the 5' mRNA sequence obtained by primer extension and dideoxysequencing. The encoded protein of 574 amino acids is extremely hydrophobic and has a molecular weight of 63371 daltons. The site of proteolytic cleavage within this protein was accurately mapped by determining a partial amino acid sequence of the N-terminus of the larger subunit (F1) purified by radioimmunoprecipitation using monoclonal antibodies. Alignment of the N-terminus of the F1 subunit within the deduced amino acid sequence of Fo permitted us to identify a sequence of lys-lys-arg-lys-arg-arg at the C-terminus of the smaller N-terminal F2 subunit that appears to represent the cleavage/activation domain. Five potential sites of glycosylation, four within the F2 subunit, were also identified. Three extremely hydrophobic domains are present in the protein; a) the N-terminal signal sequence, b) the N-terminus of the F1 subunit that is analogous to the N-terminus of the paramyxovirus F1 subunit and the HA2 subunit of influenza virus hemagglutinin, and c) the putative membrane anchorage domain near the C-terminus of F1.

Amino Acid Sequence↗

Highly recurring sequence elements identified in eukaryotic DNAs by computer analysis are often homologous to regulatory sequences or protein binding sites.

We have used computer assisted dot matrix and oligonucleotide frequency analyses to identify highly recurring sequence elements of 7-11 base pairs in eukaryotic genes and viral DNAs. Such elements are found much more frequently than expected, often with an average spacing of a few hundred base pairs. Furthermore, the most abundant repetitive elements observed in the ovalbumin locus, the beta-globin gene cluster, the metallothionein gene and the viral genomes of SV40, polyoma, Herpes simplex-1 and Mouse Mammary Tumor Virus were sequences shown previously to be protein binding sites or sequences important for regulating gene expression. These sequences were present in both exons and introns as well as promoter regions. These observations suggest that such sequences are often highly overrepresented within the specific gene segments with which they are associated. Computer analysis of other genetic units, including viral genomes and oncogenes, has identified a number of highly recurring sequence elements that could serve similar regulatory or protein-binding functions. A model for the role of such reiterated sequence elements in DNA organization and function is presented.

Animals↗

Interaction of berenil with the tyrT DNA sequence studied by footprinting and molecular modelling. Implications for the design of sequence-specific DNA recognition agents.

We have developed a technique of partially-restrained molecular mechanics enthalpy minimisation which enables the sequence-dependence of the DNA binding of a non-intercalating ligand to be studied for arbitrary sequences of considerable length (greater than = 60 base-pairs). The technique has been applied to analyse the binding of berenil to the minor groove of a 60 base-pair sequence derived from the tyrT promoter; the results are compared with those obtained by DNAse I and hydroxyl radical footprinting on the same sequence. The calculated and experimentally observed patterns of binding are in good agreement. Analysis of the modelling data highlights the importance of DNA flexibility in ligand binding. Further, the electrostatic component of the interaction tends to favour binding to AT-rich regions, whilst the van der Waals interaction energy term favours GC-rich ones. The results also suggest that an important contribution to the observed preference for binding in AT-rich regions arises from lower DNA perturbation energies and is not accompanied by reduced DNA structural perturbations in such sequences. It is therefore concluded that those modes of DNA distortion favourable to binding are probably more flexible in AT-rich regions. The structure of the modelled DNA sequence has also been analysed in terms of helical parameters. For the DNA energy-minimised in the absence of berenil, certain helical parameters show marked sequence-dependence. For example, purine-pyrimidine (R-Y) base pairs show a consistent positive buckle whereas this feature is consistently negative for Y-R pairs. Further, CG steps show lower than average values of slide while GC steps show lower than average values of rise. Similar analysis of the modelling data from the calculations including berenil highlights the importance of DNA flexibility in ligand binding. We observe that the binding of berenil induces characteristic responses in different helical parameters for the base-pairs around the binding site. For example, buckle and tilt tend to become more negative to the 5'-side of the binding site and more positive to the 3'-side, while the base steps at either side of the centre of the site show increased twist and decreased roll.

Amidines↗

Comparison of the nucleoside sequence of trpA and sequences immediately beyond the trp operon of Klebsiella aerogenes. Salmonella typhimurium and Escherichia coli.

The nucleotide sequence of trpA of Klebsiella aerogenes is presented and compared with the trpA sequences of Salmonella typhimurium and Escherichia coli. The majority of the approximately 200 differences between each pair of trpA's are single nucleotide pair changes that do not alter the amino acid sequence. Codon usage conforms to the general patterns revealed by examination of other prokaryotic gene sequences. However, codon usage in K. aerogenes trpA reflects the high G+C content of the genome of this organism. The DNA sequences just beyond trpA, the presumed transcription termination region, are also compared for the three species. Perusal of these sequences indicates that the secondary structure of the transcript segment just beyond trpA has been preserved, while the primary sequence has diverged appreciably.

Amino Acid Sequence↗

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequences representing genomic data, transcripts and proteins. Although the goal is to provide a comprehensive dataset representing the complete sequence information for any given species, the database pragmatically includes sequence data that are currently publicly available in the archival databases. The database incorporates data from over 2400 organisms and includes over one million proteins representing significant taxonomic diversity spanning prokaryotes, eukaryotes and viruses. Nucleotide and protein sequences are explicitly linked, and the sequences are linked to other resources including the NCBI Map Viewer and Gene. Sequences are annotated to include coding regions, conserved domains, variation, references, names, database cross-references, and other features using a combined approach of collaboration and other input from the scientific community, automated annotation, propagation from GenBank and curation by NCBI staff.

Animals↗

How many clones need to be sequenced from a single forensic or ancient DNA sample in order to determine a reliable consensus sequence?

Forensic and ancient DNA (aDNA) extracts are mixtures of endogenous aDNA, existing in more or less damaged state, and contaminant DNA. To obtain the true aDNA sequence, it is not sufficient to generate a single direct sequence of the mixture, even where the authentic aDNA is the most abundant (e.g. 25% or more) in the component mixture. Only bacterial cloning can elucidate the components of this mixture. We calculate the number of clones that need to be sampled (for various mixture ratios) in order to be confident (at various levels of confidence) to have identified the major component. We demonstrate that to be >95% confident of identifying the most abundant sequence present at 70% in the ancient sample, 20 clones must be sampled. We make recommendations and offer a free-access web-based program, which constructs the most reliable consensus sequence from the user's input clone sequences and analyses the confidence limits for each nucleotide position and for the whole consensus sequence. Accepted authentication methods must be employed in order to assess the authenticity and endogeneity of the resulting consensus sequences (e.g. quantification and replication by another laboratory, blind testing, amelogenin sex versus morphological sex, the effective use of controls, etc.) and determine whether they are indeed aDNA.

Cloning, Molecular↗

Nucleotide sequences of the trailer, nucleocapsid protein gene and intergenic regions of Newcastle disease virus strain Beaudette C and completion of the entire genome sequence.

The nucleotide sequences of the nucleocapsid protein (NP) gene, the intergenic regions in the nucleocapsid protein (NP)-phosphoprotein (P), P-matrix protein (M) and M-fusion glycoprotein gene junctions and the trailer region of a virulent Newcastle disease virus (NDV) strain Beaudette C were determined. The NP gene is 1747 nt long and encodes a protein of 489 amino acids. Each of the intergenic sequences determined is 1 nt long and, including the previously published intergenic sequences, the gene junction sequences varied in length from 1-47 nt and lacked any sequence identity. The 5' trailer region is 113 nt in length. Comparison of the sequences of the terminal leader and trailer regions of Beaudette C strain with those of nonvirulent strain B1 showed a high level of conservation, indicating the likelihood of these elements not being a factor in virulence. Together with previously published data, this report completes the sequence of the 15,186 nt genomic RNA of NDV strain Beaudette C.

Base Sequence↗

Nucleotide sequence of the pbpA gene and characteristics of the deduced amino acid sequence of penicillin-binding protein 2 of Escherichia coli K12.

We have determined the nucleotide sequence of the pbpA gene encoding penicillin-binding protein (PBP) 2 of Escherichia coli. The coding region for PBP 2 was 1899 base pairs in length and was preceded by a possible promoter sequence and two open reading frames. The primary structure of PBP 2, deduced from the nucleotide sequence, comprised 633 amino acid residues. The relative molecular mass was calculated to be 70867. The deduced sequence agreed with the NH2-terminal sequence of PBP 2 purified from membranes, suggesting that PBP 2 has no signal peptide. The hydropathy profile suggested that the NH2-terminal hydrophobic region (a stretch of 25 non-ionic amino acids) may anchor PBP 2 in the cytoplasmic membrane as an ectoprotein. There were nine homologous segments in the amino acid sequence of PBP 2 when compared with PBP 3 of E. coli. The active-site serine residue of PBP 2 was predicted to be Ser-330. Around this putative active-site serine residue was found the conserved sequence of Ser-Xaa-Xaa-Lys, which has been identified in all of the other E. coli PBPs so far studied (PBPs 1A, 1B, 3, 5 and 6) and class A and class C beta-lactamases. In the higher-molecular-mass PBPs 1A, 1B, 2 and 3, Ser-Xaa-Xaa-Lys-Pro was conserved. In the putative peptidoglycan transpeptidase domain there were six amino acid residues, which are common only in the PBPs of higher molecular mass.

Acyltransferases↗

Random AT library: autonomously replicating sequence (ARS) activity of chemically synthesized random sequences for transformation of nonconventional yeast species.

In a search for sequences that confer on bacterial plasmids the capacity of autonomous replication in yeast cells, we chemically synthesized polynucleotides 80 bp in length from an equimolar mixture of A and T. The random AT-polymer population, W80, was inserted into the plasmid YIp5-Kan1 (which carries the markers URA3 and G418(R), but does not replicate in yeast) and amplified in Escherichia coli. This library, representing 10 000 different AT sequences, was transformed into three species of yeast: Saccharomyces cerevisiae, Kluyveromyces lactis and Torulaspora delbrueckii. The aim was to evaluate the frequency, if any, of autonomously replicating sequences (ARSs) in the random sequences. A large number of transformants were obtained from each species. Many of them showed a stable transformed phenotype. Several W80 sequences were found many times for a given species, suggesting that each species preferred particular sequences for ARS function, although they are diverse in their primary sequence. In view of the high frequency and stability of the replicative plasmids found in the different hosts, this small random AT library may be conveniently used as a source of replicative gene vectors for genetic manipulation of many nonconventional yeast species, in place of searching for species-specific chromosomal ARSs.

Ascomycota↗

Polyoma virus DNA: Sequence from the late region that specifies the leader sequence for late mRNA and codes for VP2, VP3, and the N-terminus of VP1.

The DNA sequence of part of the late region of the polyoma virus genome is presented. This sequence of 1,348 nucleotide pairs encompasses the leader region for late mRNA and the coding sequence for the two minor capsid proteins VP2 and VP3. The coding sequence for the N-terminus of the major capsid protein overlaps the C-terminus of VP2/VP3 by 32 nucleotide pairs. From the DNA sequence the sizes and sequences of VP2 and VP3 could be predicted. Potential splicing signals for the processing of late mRNA's could be identified. Comparisons are made between the sequence of polyoma virus DNA and corresponding regions of simian virus 40 DNA.

Amino Acid Sequence↗

Nucleotide sequence analysis of the long terminal repeat of avian myeloblastosis virus and adjacent host sequences.

The nucleotide sequence of the integrated avian myeloblastosis virus long terminal repeat has been determined. The sequence is 385 base pairs long and is present at both ends of the viral DNA. The cell-virus junctions at each end consist of a 6-base-pair direct repeat of cell DNA next to the inverted repeat of viral DNA. The long terminal repeat also contains promoter-like sequences, an mRNA capping site, and polyadenylation signals. Several features of this long terminal repeat suggest a structural and functional similarity with sequences of transposable and other genetic elements. Comparison of these sequences with long terminal repeats of other avian retroviruses indicates that there is a great variation in the 3' unique sequence (U3), whereas the 5' specific sequences (U5) and the R region are highly conserved.

Avian Leukosis Virus↗