Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Sequence organisation in nuclear DNA from Physarum polycephalum. Genomic organisation of DNA segments containing foldback sequences.

DNA clones containing foldback sequences, derived from Physarum polycephalum nuclear DNA, can be classified according to their pattern of hybridisation to Southern blots of genomic DNA. One group of DNA clones map to unique DNA loci when used as a probe to restriction digests of Physarum nuclear DNA. These cloned segments appear to contain dispersed repetitive sequence elements located at many hundreds of sites in the genome. Similar patterns of hybridisation are generated when these cloned DNA probes are annealed to DNA restriction fragments of genomic DNA obtained from a number of different Physarum strains, indicating that no detectable alteration has occurred at these genomic loci subsequent to the divergence of the strains as a result of the introduction or deletion of mobile genetic elements. However, deletion of segments of some cloned DNA fragments occurs following their propagation in Escherichia coli. A second, distinct group of clones are shown to be derived from highly methylated segments of Physarum DNA which contain very abundant repetitive sequences with regular, though complex, arrangements of restriction sites at their various genomic locations. It is suggested that these DNA segments contain clustered repetitive sequence elements. The results lead to the conclusion that foldback elements in Physarum DNA are located in segments of the genome which display markedly different patterns of sequence organisation and degree of DNA methylation.

Base Sequence↗

Nucleotide sequence and deduced amino acid sequence of Escherichia coli adenine phosphoribosyltransferase and comparison with other analogous enzymes.

The Escherichia coli apt gene has been analyzed and its nucleotide (nt) sequence and the deduced amino acid (aa) sequence compared to those of other phosphoribosyltransferases (PRTs). The apt mRNA has a 102-nt leader sequence which may form alternate secondary structures. The RNA transcript may also form several 3' hairpin structures, which, however, do not appear to act as Rho-independent terminators. All PRTs, including E. coli adenine PRT (APRT), have a strongly conserved 13-aa sequence, as well as other regions of aa sequence or structural similarity. E. coli APRT is remarkably similar to the mouse enzyme.

Adenine Phosphoribosyltransferase↗

Rapid detection and sequencing of specific in vitro amplified DNA sequences using solid phase methods.

We describe a rapid solid phase assay for detection and sequencing of DNA sequences based on selective introduction of biotin and isotope into the specific DNA fragment amplified by the polymerase chain reaction (PCR). A two-step PCR procedure is used to lower the background signal. The in vitro amplified material is immobilized on magnetic beads with covalently coupled streptavidin and the amount of bound label is measured. Samples identified as positive can be analysed by direct solid phase DNA sequencing. A strategy is also described to use general primers for detection, capturing and sequencing, which are not homologous to the specific sequence to be detected. The concept has been optimized using oligonucleotides specific for Staphylococci and Streptococci, respectively. Here, we show that the assay can be used for detection of Plasmodium falciparum in clinical samples.

Animals↗

Comparative analysis of invertebrate Tc6 sequences that resemble the vertebrate V(D)J recombination signal sequences (RSS).

Invertebrate cells lack the p53 recombination checkpoint but contain mobile DNA sequences that transpose by a mechanism in part shared with excision of the V(D)J recombination signal sequences (RSS). In this work, inversion, deletion, and duplication of sequences associated with an invertebrate C. elegans Tc6 element is described. The structure of this C. elegans sequence and other dispersed Tc6 elements suggests that covalently closed 'hairpin' structures are not unique to excision of the V(D)J RSS by the RAG proteins, but rather can be generated by transposases at transposon termini leading to characteristic inversion and duplication events. Comparative analysis of recombination events at invertebrate sequences resembling the vertebrate V(D)J RSS may be useful in understanding V(D)J recombination-mediated recombination events in malignant vertebrate cells or genetic diseases such as ataxia telangectasia, in which the p53 recombination checkpoint is defective.

Animals↗

Genome size and the accumulation of simple sequence repeats: implications of new data from genome sequencing projects.

The relationship between the level of repetitiveness in genomic sequences and genome size has been re-investigated making use of the rapidly growing database of complete eubacterial and archaeal genome sequences combined with the fragmentary but now large amount of data from eukaryotic genomes. Relative simplicity factors (RSFs), which measure the repetitiveness of sequences, were calculated and significantly simple motifs (SSMs), which identify the kinds of sequences that are repeated, were identified. A previously reported correlation between genome size and repetitiveness was confirmed, but it was shown that the higher RSFs seen in eukaryotic genomes also reflect a generally higher level of repetitiveness independent of genome size differences. Differences in genome size are responsible for about 10% of the variance in RSF seen between species. The spectrum of SSMs seen within a genome differed markedly within the eubacteria but less so in eukaryotes and, particularly, in archaea. Species with SSM spectra that differ from the norm tend also to have high RSFs for their genome size and to be pathogens that make use of repetitive sequences to avoid host defence responses. Some of the variance in repetitiveness seen in other species may therefore also reflect the action of selection, although other forces such as variation in the effectiveness of mechanisms for regulating slippage errors of replication, may also be important.

Animals↗

Automated SNP detection in expressed sequence tags: statistical considerations and application to maritime pine sequences.

We developed an automated pipeline for the detection of single nucleotide polymorphisms (SNPs) in expressed sequence tag (EST) data sets, by combining three DNA sequence analysis programs: Phred, Phrap and PolyBayes. This application requires access to the individual electrophoregram traces. First, a reference set of 65 SNPs was obtained from the sequencing of 30 gametes in 13 maritime pine (Pinus pinaster Ait.) gene fragments (6671 bp), resulting in a frequency of 1 SNP every 102.6 bp. Second, parameters of the three programs were optimized in order to retrieve as many true SNPs, while keeping the rate of false positive as low as possible. Overall, the efficiency of detection of true SNPs was 83.1%. However, this rate varied largely as a function of the rare SNP allele frequency: down to 41% for rare SNP alleles (frequency < 10%), up to 98% for allele frequencies above 10%. Third, the detection method was applied to the 18498 assembled maritime pine (Pinus pinaster Ait.) ESTs, allowing to identify a total of 1400 candidate SNPs, in contigs containing between 4 and 20 sequence reads. These genetic resources, described for the first time in a forest tree species, were made available at http://www.pierroton.inra/genetics/Pinesnps. We also derived an analytical expression for the SNP detection probability as a function of the SNP allele frequency, the number of haploid genomes used to generate the EST sequence database, and the sample size of the contigs considered for SNP detection. The frequency of the SNP allele was shown to be the main factor influencing the probability of SNP detection.

Algorithms↗

The genome sequence of the food-borne pathogen Campylobacter jejuni reveals hypervariable sequences.

Campylobacter jejuni, from the delta-epsilon group of proteobacteria, is a microaerophilic, Gram-negative, flagellate, spiral bacterium-properties it shares with the related gastric pathogen Helicobacter pylori. It is the leading cause of bacterial food-borne diarrhoeal disease throughout the world. In addition, infection with C. jejuni is the most frequent antecedent to a form of neuromuscular paralysis known as Guillain-Barré syndrome. Here we report the genome sequence of C. jejuni NCTC11168. C. jejuni has a circular chromosome of 1,641,481 base pairs (30.6% G+C) which is predicted to encode 1,654 proteins and 54 stable RNA species. The genome is unusual in that there are virtually no insertion sequences or phage-associated sequences and very few repeat sequences. One of the most striking findings in the genome was the presence of hypervariable sequences. These short homopolymeric runs of nucleotides were commonly found in genes encoding the biosynthesis or modification of surface structures, or in closely linked genes of unknown function. The apparently high rate of variation of these homopolymeric tracts may be important in the survival strategy of C. jejuni.

Amino Acid Sequence↗

The partial amino acid sequence of bovine cartilage proteoglycan, deduced from a cDNA clone, contains numerous Ser-Gly sequences arranged in homologous repeats.

We have determined the sequence of a partial cDNA clone encoding the C-terminal region of bovine cartilage aggregating proteoglycan core protein. The deduced amino acid sequence contains a cysteine-rich region which is homologous with chicken hepatic lectin. This lectin-homologous region has previously been identified in rat and chicken cartilage proteoglycan. The bovine sequence presented here is highly homologous with the rat and chicken amino acid sequences in this apparently globular region. A region containing clusters of Ser-Gly sequences is located N-terminal to the lectin homology domain. These Ser-Gly-rich segments are arranged in tandemly repeated, approx. 100-residue-long, homology domains. Each homology domain consists of an approx. 75-residue-long Ser-Gly-rich region separated by an approx. 25-residue-long segment lacking Ser-Gly dipeptides. These dipeptides are arranged in 10-residue-long segments in the 100-residue-long homology domains. The shorter homologous segments are tandemly repeated some six times in each 100-residue-long homology domain. Serine residues in these repeats are potential attachment sites for chondroitin sulphate chains.

Amino Acid Sequence↗

A systematic analysis of sequences of human antiphospholipid and anti-beta2-glycoprotein I antibodies: the importance of somatic mutations and certain sequence motifs.

OBJECTIVE: Previous studies have suggested the importance of somatic mutations and certain residues in the complementarity determining regions (CDRs) of antiphospholipid antibodies (aPL) implicated in the pathogenesis of antiphospholipid antibody syndrome (APS). The authors tested this hypothesis by carrying out a systematic analysis of all published aPL sequences. METHODS: Each aPL variable region sequence was aligned to the closest germline counterpart in the VBASE Sequence Directory by using DNAPLOT software, allowing analysis of nucleotide homology and distribution of somatic mutations. The probability that this distribution arose as a result of antigen-driven accumulation of replacement mutations in the CDRs was tested statistically. RESULTS: There was no preferential gene or family use in the 36 aPL sequences identified. Immunoglobulin (Ig) M aPL had few somatic mutations compared with IgG. Of the IgG aPL, 9 of 14 showed evidence of antigen-driven accumulation of replacement mutations in the CDRs. Multinomial analysis allowed a clearer statistical identification of sequences that had been subject to antigen drive. The more specific IgM aPL and some IgG aPL displayed an accumulation of arginine, asparagine, and lysine residues in CDRs. CONCLUSIONS: High-specificity binding in IgG aPL, but not in more specific IgM aPL, is conferred by antigen-driven somatic mutation. This may in part be caused by an accumulation of arginine, asparagine, and lysine residues in the CDRs, which are germlines encoded in the more specific IgM aPL, but often arise because of somatic mutation in IgG aPL. RELEVANCE: An understanding of the role of arginine, asparagine, and lysine residues in the binding of pathogenic aPL to phospholipids, and to beta(2)-glycoprotein I, may eventually help in the development of drugs to interfere with those interactions, and thereby improve the treatment of antiphospholipid antibody syndrome.

Amino Acid Sequence↗

Nucleotide sequence divergence of mouse immunoglobulin gamma 1 and gamma 2b chain genes and the hypothesis of intervening sequence-mediated domain transfer.

The nucleotide sequences of the constant portions of mouse immunoglobulin gamma 1 and gamma 2b genes were compared. A remarkable homology was found in a long (about 500 nucleotides) continuous segment including the entire CH1 coding region and about the first half of the first intervening sequence. Furthermore, comparison of amino acid sequences of four gamma-class chains revealed that the CH1 domain shows limited divergence among gamma 1, gamma 2a, and gamma 2b. Interestingly, the homology region extends to the CH2 domain in gamma 2a and gamma 2b. These findings suggest that, during their evolution, a double unequal crossing-over event has taken place at different intervening sequences, resulting in the transfer of the DNA segment coding for the CH1 domain or CH1-CH2 domains. A possible evolutionary implication for such an "intervening sequence-mediated domain transfer" event is discussed.

Amino Acid Sequence↗

Rat pancreatic kallikrein mRNA: nucleotide sequence and amino acid sequence of the encoded preproenzyme.

We have cloned via recombinant DNA technology the mRNA sequence for rat pancreatic preprokallikrein. Four cloned overlapping double-stranded cDNAs gave a continuous mRNA sequence of 867 nucleotides beginning within the 5'-noncoding region and extending to the poly(A) tail. The mRNA sequence reveals that pancreatic kallikrein is synthesized as a prezymogen of 265 amino acids, including a proposed secretory prepeptide of 17 amino acids and a proposed activation peptide of 11 amino acids. The activation peptide, although similar in length, is distinct from those of the other classes of pancreatic serine proteases. The amino acid sequence of the predicted active form of the enzyme is closely related to the partial sequences obtained for other kallikrein-like serine proteases including rat submaxillary gland kallikrein, pig pancreatic and submaxillary gland kallikreins, the gamma subunit of mouse nerve growth factor, and rat tonin. Key amino acid residues thought to be involved in the substrate-cleavage specificity of kallikreins are retained. Hybridization analysis showed relatively high levels of kallikrein mRNA in the rat pancreas, submaxillary and parotid glands, spleen, and kidney, indicating the active synthesis of kallikrein in these tissues.

Amino Acid Sequence↗

The complete cDNA and deduced amino acid sequence of a type II mouse epidermal keratin of 60,000 Da: analysis of sequence differences between type I and type II keratins.

We present the complete nucleotide and deduced amino acid sequences of a mouse epidermal keratin subunit of 60,000 Da. The keratin possesses a central alpha-helical domain of four tracts (termed 1A, 1B, 2A, and 2B) that can form coiled-coils, interspersed by short linker sequences, and has non-alpha-helical terminal domains. This pattern of secondary structure is emerging as common to all intermediate filament subunits. The alpha-helical sequences conform to the type II class of keratins. Accordingly, this is the first type II keratin for which complete sequence information is available, and thus it facilitates elucidation of the fundamental distinctions between type I and type II keratins. It has been observed that type I keratins are acidic and type II keratins are neutral--basic in charge. We suggest that the basis for this empirical correlation between type and charge resides in the respective net charges of the 1A and 2B tracts. Calculations on interchain interactions between charged residues in the alpha-helical domains indicate that this keratin prefers to participate in dimers according to an in-register parallel arrangement. The terminal domains of this keratin possess characteristic glycine-rich sequences, and the carboxyl-terminal domain is highly homologous to that of a human epidermal keratin of 56,000 Da. According to the hypothesis that end-domains are located on the periphery of keratin filaments, we conclude that the corresponding mouse and human keratins are closely related, both structurally and functionally.

Amino Acid Sequence↗

Molecular cloning and primary nucleotide sequence analysis of a distinct human immunodeficiency virus isolate reveal significant divergence in its genomic sequences.

In an effort to evaluate data on genomic relatedness among the various human immunodeficiency viruses (HIVs), we have molecularly cloned a virus isolate designated HIV (CDC-451). Preliminary characterization of the HIV (CDC-451) clone indicated that the restriction enzyme map was distinct from those of other known HIV isolates. Analysis of the primary nucleotide sequence of the regions encoding the structural proteins and comparison with sequences known for other HIV isolates indicated substantial differences for HIV (CDC-451). The sequences encoding the group-specific antigen gene, although they showed some variation, were conserved to a greater extent than were those encoding envelope proteins. In the envelope gene sequences, most of the changes (up to 24.5% divergence) were located in the amino-terminal region encoding a glycoprotein with a Mr of 120,000. The carboxyl-terminal region, encoding a protein of Mr 41,000, was more highly conserved. The variation in the sequences encoding envelope proteins may have important implications for the antigenic properties and/or pathogenicity of the disease and for its detection and ultimate eradication.

Adolescent↗

A DNA sequencing strategy that requires only five bases of known terminal sequence for priming.

We have previously reported an enhanced version of sequencing by hybridization (SBH), termed positional SBH (PSBH). PSBH uses partially duplex probes containing single-stranded 3' overhangs, instead of simple single-stranded probes. Stacking interactions between the duplex probe and a single-stranded target allow us to reduce the probe sizes required to 5-base single-stranded overhangs. Here we demonstrate the use of PSBH to capture relatively long single-stranded DNA targets and perform standard solid-state Sanger sequencing on these primer-template complexes without ligation. Our results indicate that only 5 bases of known terminal sequence are required for priming. In addition, the partially duplex probes have the ability to capture their specific target from a mixture of five single-stranded targets with different 3'-terminal sequences. This indicates the potential utility of the PSBH approach to sequence mixtures of DNA targets without prior purification.

Base Sequence↗

Human cholesteryl ester transfer protein gene proximal promoter contains dietary cholesterol positive responsive elements and mediates expression in small intestine and periphery while predominant liver and spleen expression is controlled by 5'-distal sequences. Cis-acting sequences mapped in transgenic mice.

The plasma cholesteryl ester transfer protein (CETP) facilitates the transfer of high density lipoprotein cholesteryl esters to other lipoproteins and appears to be a key regulated component of reverse cholesterol transport. Earlier studies showed that a CETP transgene containing natural flanking sequences (-3.4 kilobase pairs (kbp) upstream, +2.2 kbp downstream) was expressed in an authentic tissue distribution and induced in liver and other tissues in response to dietary or endogenous hypercholesterolemia. In order to localize the DNA elements responsible for these effects, we prepared transgenic mice expressing six new DNA constructs containing different amounts of natural flanking sequence of the CETP gene. Tissue-specific expression and dietary cholesterol response of CETP mRNA were determined. The native pattern of predominant expression in liver and spleen with cholesterol induction was shown by a -3.4 (5'), +0.2 (3') kbp transgene, indicating no major contribution of distal 3'-sequences. Serial 5'-deletions showed that a -570 base pairs (bp) transgene gave predominant expression in small intestine with cholesterol induction of CETP mRNA in that organ, and a -370 bp transgene gave highest expression in adrenal gland with partial dietary cholesterol induction of CETP mRNA and plasma activity. Further deletion to -138 bp 5'-flanking sequence resulted in a transgene that was not expressed in vivo. Both the -3.4 kbp and -138 bp transgenes were expressed when transfected into a cultured murine hepatocyte cell line, but only the former was induced by treating the cells with LDL. When linked to a human apoA-I transgene, the -570 to -138 segment of the CETP gene promoter gave rise to a relative positive response of hepatic apoA-I mRNA to the high cholesterol diet in two out of three transgenic lines. Thus, 5'-elements between -3,400 and -570 bp in the CETP promoter endow predominant expression in liver and spleen. Elements between -570 and -370 are required for expression in small intestine and some other tissues, and elements between -370 and -138 contribute to adrenal expression. The minimal CETP promoter element associated with a positive sterol response in vivo was found in the proximal CETP gene promoter between -370 and -138 bp. This region contains a tandem repeat of a sequence known to mediate sterol down-regulation of the HMG-CoA reductase gene, suggesting either the presence of separate positive and negative sterol response elements in this region or the use of a common DNA element for both positive and negative sterol responses.

Animals↗

Prediction of the coding sequences of unidentified human genes. XVII. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

To provide information regarding the coding sequences of unidentified human genes, we have conducted a sequencing project of human cDNAs which encode large proteins. We herein present the entire sequences of 100 cDNA clones of unknown human genes, named KIAA1444 to KIAA1543, from two sets of size-fractionated human adult and fetal brain cDNA libraries. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here were 4.4 kb and 2.6 kb (856 amino acid residues), respectively. Database searches of the predicted amino acid sequences classified 53 predicted gene products into the following five functional categories: cell signaling/communication, nucleic acid management, cell structure/motility, protein management and metabolism. It was also revealed that homologues for 32 KIAA gene products were detected in the databases, which were similar in sequence through almost their entire regions. Additionally, the chromosomal loci of the genes were determined by using human-rodent hybrid panels unless their chromosomal loci were already assigned in the public databases. The expression levels of the genes were monitored in spinal cord, fetal brain and fetal liver, as well as in 10 human tissues and 8 brain regions, by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

Invariant chains with the class II binding site replaced by a sequence from influenza virus matrix protein constrain low-affinity sequences to MHC II presentation.

Presentation of antigenic peptides by MHC II molecules is required to initiate CD4 T(h) cell responses. Some peptides, however, because of low affinity for MHC II, are not efficiently presented. A segment of the MHC II chaperon molecule, invariant chain (Ii), is known to bind early in biosynthesis with low affinity to the peptide binding groove. Here we have exploited the properties of Ii to manipulate the MHC II-loading pathway and to present low-affinity sequences. We used a deletion mutant of Ii where the promiscuous binding site to MHC II, which is adjacent to the groove binding segment, was deleted. A recombinant Ii (rIi) chimera, derived from this construct, was made in which the class II binding segment was exchanged for wild-type or single amino acid substitution variants of an HLA-DR1-restricted sequence from influenza matrix protein (MAT), which leads to MHC II allotype-specific binding. This rIi was expressed in antigen-presenting cells (APC) and introduced the MAT sequence into the MHC II-processing pathway. As expected, rIiMAT elicited antigen-specific, DR1-restricted T cell cytokine production and proliferation. Significantly, rIiMAT, that binds the HLA-DR4 allele with low affinity, elicited DR4-restricted IL-2 production but not proliferation. In contrast, exogenously provided MAT peptide failed to elicit any responses from DR4-restricted T cells. Compatible results were obtained with a single amino acid substitution variant (MAT(T)), which binds with high affinity to DR4 but low affinity to DR1. We conclude that loading of MHC II with antigenic peptides from endogenously synthesized rIi chimeras allows presentation of low-affinity sequences that cannot be presented if provided exogenously as peptides. Ii fusion proteins containing low-affinity antigenic sequences might be useful for vaccination with tumor antigens to overcome deficiencies in antigen presentation.

Adult↗

Microcomputer programs for back translation of protein to DNA sequences and analysis of ambiguous DNA sequences.

Three computer programs are described which may be used to translate a DNA sequence into a protein sequence, back translate the protein sequence into an ambiguous DNA sequence, and then do pattern searching in the ambiguous sequence. The programs are written in the C programming language, have been compiled to run on a microcomputer under the CP/M 80 operating system, and may be copied in binary format through a modem. They are also to become available for the IBM/PC.

Amino Acid Sequence↗