Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

A 14-3-3 protein of Chlamydomonas reinhardtii associated with the endoplasmic reticulum: nucleotide sequence of the cDNA and the corresponding gene and derived amino acid sequence.

Two major 14-3-3 proteins of the unicellular green alga Chlamydomonas reinhardtii were purified and partially sequenced. The obtained data show that the 30-kDa isoform predominant in the cytosol is encoded by a previously cloned and sequenced 14-3-3 cDNA whereas the 27-kDa isoform represents a new 14-3-3 protein which is largely associated with the endoplasmic reticulum (ER). Therefore, the corresponding cDNA was cloned and sequenced. The nucleotide sequence of this new cDNA species and the derived amino acid sequence differ considerably from the previously cloned Chlamydomonas 14-3-3 cDNA. The conclusion that the divergent evolution of the corresponding genes must have started rather early as compared to the 14-3-3 genes of other organisms was corroborated by their different genomic organization. The amino acid sequences of both 14-3-3 isoforms were comparatively analysed to find differences which might be responsible for their differential binding to the ER.

14-3-3 Proteins↗

Inter-specific sequence conservation and intra-individual sequence variation in a spider silk gene.

Currently, studies on major ampullate spidroin 1 (MaSp1) genes of non-orb weaving spiders are few, and it is not clear whether genes of these organisms exhibit the same characteristics as those of orb-weavers. In addition, many studies have proposed that MaSp1 might be a single gene with allelic variants, but supporting evidence is still lacking. In this study, we compared partial DNA and amino acid sequences of MaSp1 cloned from different spider guilds. We also cloned partial MaSp1 sequences from genomic DNA and cDNA of the same individuals of spiders using the same primer combination to see if different molecular forms existed. In the repetitive region of partial MaSp1 sequences obtained, GGX, GA and poly-A motifs were present in all Araneomorphae and Mygalomorpae species examined. An extreme similarity in MaSp1 non-repetitive portions was found in sequences of ecribellate, cribellate and Mygalomorphae web-builders and such a result suggested that this sequence might exhibit an important function. A comparison of sequences amplified from the same individual showed that substitutions in amino acids occurred in both repetitive and non-repetitive regions, with a much higher variation in the former. These results suggest that the MaSp1 of Araneomorphae spiders exhibits several forms in an individual spider and it might be either a multiple gene or a single gene with a multiple exon/intron organization.

Amino Acid Sequence↗

Sequence similarities between large subunit ribosomal RNA gene intervening sequences from different Helicobacter species.

When the 23S rRNA genes from several Helicobacter species were amplified by PCR and compared with similar amplicons derived from H. pylori, they were seen to be enlarged in size. Sequencing of these enlarged genes from H. mustelae, H. canis (two strains) and H. muridarum identified insertions of novel sequence (intervening sequences, IVSs) sized between 93 and 377 bp located at nt 545, in place of an 8-nt sequence in the conventionally sized H. pylori gene. These IVSs were not present elsewhere in the genome. All strains with such IVSs lacked intact 23S rRNA which was replaced by two fragment whose sizes were consistent with cleavage at either side of the particular IVS. The predicted secondary structures of the four IVSs were characterised by base pairing at the 5' and 3' ends to form a stem. The four IVSs exhibited significant sequence inter-relationships. Further relationships were also observed between them and similar elements in both small and large subunit rRNA genes of other Helicobacter and Campylobacter species. Alignment of each IVS with the other such elements identified blocks of related sequence consistent with insertion/deletion events, indicating possible evolutionary relationships.

Base Sequence↗

Amino acid sequence of the BSC-1 cell growth inhibitor (polyergin) deduced from the nucleotide sequence of the cDNA.

The complete amino acid sequence of the BSC-1 cell growth inhibitor, including its precursor polypeptide, is reported. The sequence was deduced from the nucleotide sequence of the cDNA. The N-terminal amino acid sequence of the mature bioactive BSC-1 cell growth inhibitor is identical with the N-terminal sequences of the factors that have been called type beta 2 transforming growth factor and cartilage-inducing factor B, suggesting that these are identical. The complete amino acid sequence of the mature BSC-1 cell growth inhibitor differs from that of human type beta transforming growth factor in 32 of the 112 amino acids. Polyergin is proposed as the name for the BSC-1 cell growth inhibitor.

Amino Acid Sequence↗

Sequence of the human 40-kDa keratin reveals an unusual structure with very high sequence identity to the corresponding bovine keratin.

The complete amino acid and DNA sequences of the human 40-kDa keratin are reported. The DNA sequence encodes a protein of 44,098 Da, which is unique in that it lacks the terminal non-alpha-helical tail segment found in all other keratins. When the human 40-kDa keratin amino acid sequence is compared to the corresponding bovine keratin, the overall identity is 89%. The coil-forming regions are 89% identical and the head regions are 88% identical. This similarity is also evident in the DNA sequence of the coding region, the 5' upstream sequences, and the 3' noncoding sequences. The high degree of cross-species identity between bovine and human 40-kDa keratins suggests that there is strong evolutionary pressure to conserve the structure of this keratin. This in turn suggests an important and universal role for this intermediate filament subunit in all species.

Amino Acid Sequence↗

Relationship between P-box amino acid sequence and DNA binding specificity of the thyroid hormone receptor. The effects of half-site sequence in everted repeats.

The three P-box amino acids in the DNA recognition alpha-helix of steroid/thyroid hormone receptors participate in the discrimination of the central base pairs of the hexameric half-sites of receptor response elements in DNA. A series of 57 variants of the beta isoform of the human thyroid hormone receptor were constructed in which the 19 possible amino acid substitutions were incorporated at each of the three P-box positions. The effects of these substitutions on the sequence specificity of the DNA binding activity of the receptor were analyzed using 16 everted repeat elements which differed in sequence in the two central base pairs of the hexameric half-sites. Only receptors with glutamate or aspartate as the first P-box amino acid had detectable DNA binding affinity on everted repeats with AGGNCA half-sites. Only those receptors with alanine, glycine, serine, or proline in the second P-box position were able to bind to this same group of everted repeat elements. In contrast, many of the variant receptors with substitutions at the third P-box position were capable of binding to the AGGNCA group of repeat elements. The actual substitutions at the third P-box position that were compatible with binding depended upon the identity of the fourth base pair of the AGGNCA half-sites. Of the remaining 12 everted repeat sequences, only those with AGTTCA or AGTCCA half-sites were able to bind any of the receptors. In addition to wild type receptor, several variant receptors with amino acid substitutions in either the first or third P-box position were able to bind to the everted repeat with AGTTCA half-sites. The everted repeat with AGTCCA half-sites was bound by receptors with a DGG, NGG, or EGQ P-box sequence, but not the wild type receptor which has an EGG P-box sequence. These data demonstrate that all three P-box positions of the thyroid hormone receptor function to discriminate between half-sites that differ in sequence at the third and fourth base pairs.

Amino Acid Sequence↗

Methods for comparing a DNA sequence with a protein sequence.

We describe two methods for constructing an optimal global alignment of, and an optimal local alignment between, a DNA sequence and a protein sequence. The alignment model of the methods addresses the problems of frameshifts and introns in the DNA sequence. The methods require computer memory proportional to the sequence lengths, so they can rigorously process very huge sequences. The simplified versions of the methods were implemented as computer programs named NAP and LAP. The experimental results demonstrate that the programs are sensitive and powerful tools for finding genes by DNA-protein sequence homology.

Algorithms↗

Optimal alignment between groups of sequences and its application to multiple sequence alignment.

Four algorithms, A-D, were developed to align two groups of biological sequences. Algorithm A is equivalent to the conventional dynamic programming method widely used for aligning ordinary sequences, whereas algorithms B-D are designed to evaluate the cost for a deletion/insertion more accurately when internal gaps are present in either or both groups of sequences. Rigorous optimization of the 'sum of pairs' (SP) score is achieved by algorithm D, whose average performance is close to O(MNL2), where M and N are numbers of sequences included in the two groups and L is the mean length of the sequences. Algorithm B uses some approximations to cope with profile-based operations, whereas algorithm C is a simpler variant of algorithm D. These group-to-group alignment algorithms were applied to multiple sequence alignment with two iterative strategies: a progressive method based on a given binary tree and a randomized grouping--realignment method. The advantages and disadvantages of the four algorithms are discussed on the basis of the results of examinations of several protein families.

Algorithms↗

Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%.

MOTIVATION: Searching for non-coding RNA (ncRNA) genes and structural RNA elements (eleRNA) are major challenges in gene finding today as these often are conserved in structure rather than in sequence. Even though the number of available methods is growing, it is still of interest to pairwise detect two genes with low sequence similarity, where the genes are part of a larger genomic region. RESULTS: Here we present such an approach for pairwise local alignment which is based on foldalign and the Sankoff algorithm for simultaneous structural alignment of multiple sequences. We include the ability to conduct mutual scans of two sequences of arbitrary length while searching for common local structural motifs of some maximum length. This drastically reduces the complexity of the algorithm. The scoring scheme includes structural parameters corresponding to those available for free energy as well as for substitution matrices similar to RIBOSUM. The new foldalign implementation is tested on a dataset where the ncRNAs and eleRNAs have sequence similarity <40% and where the ncRNAs and eleRNAs are energetically indistinguishable from the surrounding genomic sequence context. The method is tested in two ways: (1) its ability to find the common structure between the genes only and (2) its ability to locate ncRNAs and eleRNAs in a genomic context. In case (1), it makes sense to compare with methods like Dynalign, and the performances are very similar, but foldalign is substantially faster. The structure prediction performance for a family is typically around 0.7 using Matthews correlation coefficient. In case (2), the algorithm is successful at locating RNA families with an average sensitivity of 0.8 and a positive predictive value of 0.9 using a BLAST-like hit selection scheme. AVAILABILITY: The program is available online at http://foldalign.kvl.dk/

Algorithms↗

A first full outer capsid protein sequence data-set in the Orbivirus genus (family Reoviridae): cloning, sequencing, expression and analysis of a complete set of full-length outer capsid VP2 genes of the nine African horsesickness virus serotypes.

The outer capsid protein VP2 of African horsesickness virus (AHSV) is a major protective antigen. We have cloned full-length VP2 genes from the reference strains of each of the nine AHSV serotypes. Baculovirus recombinants expressing the cloned VP2 genes of serotypes 1, 2, 4, 6, 7 and 8 were constructed, confirming that they all have full open reading frames. This work completes the cloning and expression of the first full set of AHSV VP2 genes. The clones of VP2 genes of serotypes 1, 2, 5, 7 and 8 were sequenced and their amino acid sequences were deduced. Our sequencing data, together with that of the published VP2 genes of serotypes 3, 4, 6 and 9, were used to generate the first complete sequence analysis of all the (sero)types for a species of the Orbivirus genus. Multiple alignment of the VP2 protein sequences showed that homology between all nine AHSV serotypes varied between 47.6 % and 71.4 %, indicating that VP2 is the most variable AHSV protein. Phylogenetic analysis grouped together the AHSV VP2s of serotypes that cross-react serologically. Low identity between serotypes was demonstrated for specific regions within the VP2 amino acid sequences that have been shown to be antigenic and play a role in virus neutralization. The data presented here impact on the development of new vaccines, the identification and characterization of antigenic regions, the development of more rapid molecular methods for serotype identification and the generation of comprehensive databases to support the diagnosis, epidemiology and surveillance of AHS.

African Horse Sickness↗

TFIID sequence recognition of the initiator and sequences farther downstream in Drosophila class II genes.

Immunopurified TFIID produces a large DNase I footprint over the hsp70, hsp26, and histone H3 promoters of Drosophila. These footprints span from the TATA element to a position approximately 35 nucleotides downstream from the transcription start site. Using a "missing nucleoside" analysis, four regions within the three promoters have been found to be important for TFIID binding: the TATA element, the initiator, and two regions located approximately 18 and 28 nucleotides downstream of the transcription start site. On the basis of the missing nucleoside data, the initiator appears to contribute as much to the affinity as the TATA element. However, there is weak conservation of the sequence in this region. To determine whether a preferred binding sequence exists in the vicinity of the initiator, the nucleotide composition of this region within the hsp70 promoter was randomized and then subjected to selection by TFIID. After five rounds of selection, the preferred sequence motif--G/A/T C/TAT/GTG--emerged. This motif is a close match to consensus sequences that have been derived by comparing the initiator region of numerous insect promoters. Selection of this sequence demonstrates that sequence-specific interactions downstream of the TATA element contribute to the interaction of TFIID on a wide spectrum of promoters.

Animals↗

Conserved serine-rich sequences in xylanase and cellulase from Pseudomonas fluorescens subspecies cellulosa: internal signal sequence and unusual protein processing.

The complete nucleotide sequence of the xynA gene coding for a xylanase (XYLA) expressed by Pseudomonas fluorescens subspecies cellulosa, has been determined. The structural gene consists of an open reading frame of 1833 bp followed by a TAA stop codon. Confirmation of the nucleotide sequence was obtained by comparing the predicted amino acid sequence with that derived by N-terminal analysis of purified forms of the xylanase. The signal peptide present at the N terminus of mature XYLA closely resembles signal peptides of other secreted proteins. Truncated forms of the xylanase gene, in which the sequence encoding the N-terminal signal peptide had been deleted, still expressed coli. XYLA contains domains which are homologous to an endoglucanase expressed by the same organism. These structures include serine-rich sequences. Bal31 deletions of xynA revealed the extent to which these conserved sequences, in XYLA, were essential for xylanase activity. Downstream of the TAA stop codon is a G + C-rich region of dyad symmetry (delta G = 24 kcal) characteristic of E. coli Rho-independent transcription terminators.

Amino Acid Sequence↗

HLA-C high resolution typing: analysis of exons 2 and 3 by sequence based typing and detection of polymorphisms in exons 1-5 by sequence specific primers.

HLA-C high resolution sequence based typing developed in this study involves a unique DNA amplification encompassing exon 1 to intron 3 and four fluorescent sequencing reactions covering exon 2 and 3. Both dye primer and dye terminator sequencing techniques were performed and results compared. This approach allowed the identification of all of the 50 HLA-C allelic variants so far described, except for two allele pairs that are distinguished by non-coding nucleotide changes (Cw*12021 = 12022, Cw*15051 = 15052) and three allele pairs (Cw*0701 = 706, Cw*1701 = 1702 and Cw*1801 = 1802) that share the same nucleotide sequence in exon 2 and 3. For complete subtyping of these allelic variants, an amplification based on sequence specific primers (PCR-SSP) was used. No ambiguous heterozygous combinations of alleles were detected in our panel so far. HLA-C typing data obtained by this method were compared with data from serological and low resolution PCR-SSP typing, which had been performed previously on the samples sequenced.

Base Sequence↗

Amino acid sequence of Coprinus macrorhizus peroxidase and cDNA sequence encoding Coprinus cinereus peroxidase. A new family of fungal peroxidases.

Sequence analysis and cDNA cloning of Coprinus peroxidase (CIP) were undertaken to expand the understanding of the relationships of structure, function and molecular genetics of the secretory heme peroxidases from fungi and plants. Amino acid sequencing of Coprinus macrorhizus peroxidase, and cDNA sequencing of Coprinus cinereus peroxidase showed that the mature proteins are identical in amino acid sequence, 343 residues in size and preceded by a 20-residue signal peptide. Their likely identity to peroxidase from Arthromyces ramosus is discussed. CIP has an 8-residue, glycine-rich N-terminal extension blocked with a pyroglutamate residue which is absent in other fungal peroxidases. The presence of pyroglutamate, formed by cyclization of glutamine, and the finding of a minor fraction of a variant form lacking the N-terminal residue, indicate that signal peptidase cleavage is followed by further enzymic processing. CIP is 40-45% identical in amino-acid sequence to 11 lignin peroxidases from four fungal species, and 42-43% identical to the two known Mn-peroxidases. Like these white-rot fungal peroxidases, CIP has an additional segment of approximately 40 residues at the C-terminus which is absent in plant peroxidases. Although CIP is much more similar to horseradish peroxidase (HRP C) in substrate specificity, specific activity and pH optimum than to white-rot fungal peroxidases, the sequences of CIP and HRP C showed only 18% identity. Hence, CIP qualifies as the first member of a new family of fungal peroxidases. The nine invariant residues present in all plant, fungal and bacterial heme peroxidases are also found in CIP. The present data support the hypothesis that only one chromosomal CIP gene exists. In contrast, a large number of secretory plant and fungal peroxidases are expressed from several peroxidase gene clusters. Analyses of three batches of CIP protein and of 49 CIP clones revealed the existence of only two highly similar alleles indicating less peroxidase polymorphism in C. cinereus strains than observed in plants and white-rot fungi.

Amino Acid Sequence↗

Species of tetrahymena identical by small subunit rRNA gene sequences are discriminated by mitochondrial cytochrome c oxidase I gene sequences.

The mitochondrial cytochrome c oxidase 1 (CO1) genes of two isolates of each of the seven mating types of Tetrahymena thermophila were sequenced and found to differ by < 1% in nucleotide sequence and to be identical by putative protein sequence. As this gene was highly conserved in this species, the CO1 gene sequence was determined for four pairs of Tetrahymena species identical in their small subunit rRNA gene sequences. The following pairs of species showed from 1% to 12% divergence at the nucleotide level, enabling discrimination of all these species: (1) Tetrahymena pyriformis strain T and Tetrahymena setosa strain HZ-1; (2) Tetrahymena canadensis strain UM1215 and Tetrahymena rostrata strain ID-3; (3) Tetrahymena pigmentosa strain UM1285 and Tetrahymena hyperangularis strain EN112; and (4) Tetrahymena tropicalis strain TC-105 and Tetrahymena mobilis. However, because of the synonymous nature of the majority of substitutions, the pairs of species were identical based on the putative protein sequence.

Animals↗

Allele Level Sequencing of Killer Cell Immunoglobulin-Like Receptor Genes Using Oxford Nanopore Long Read Sequencing.

The human Killer cell Immunoglobulin-like Receptor (KIR) genes, found on chromosome 19, encode for cell surface protein receptors that, through interaction with their ligand, modulate the action of Natural Killer (NK) cells and some subsets of T lymphocytes. KIR genes exhibit extensive variation through variable gene content, copy number, and allele polymorphism. The combination of KIR genes and their ligands is implicated in various clinical settings including haematopoietic stem cell and solid organ transplant, and infectious disease progression. KIR gene content has been used in the selection of optimal stem cell donors with haplotype variations in recipient and donor giving differential clinical outcomes. With the introduction of massively parallel clonal next generation sequencing and single molecule long read third generation sequencing, allele level determination of KIR genotypes has become feasible. We describe a method for amplicon-based long read sequencing on the Oxford Nanopore Technologies platform that provides largely unambiguous allele level typing of KIR genes. The method was validated using DNA extracted from 48 10th International Histocompatibility Workshop (IHWS) cell lines with previously published allele level KIR genotypes and 176 Western Australian samples previously tested for the presence or absence of KIR genes. Our long-read sequencing method was able to accurately determine KIR alleles with an overall concordance of 97%-99% with the published data. Importantly, phasing ambiguity caused by the inability to phase heterozygous base positions over long stretches of gene sequence was resolved in several samples. Thus, our long read PCR sequencing strategy can be used to determine KIR genotypes at allele resolution level.

Humans↗

Flagellar transcriptional activators FlbB and FlaI: gene sequences and 5' consensus sequences of operons under FlbB and FlaI control.

The regulation of the expression of the operons in the flagella-chemotaxis regulon in Escherichia coli has been shown to be a highly ordered cascade which closely parallels the assembly of the flagellar structure and the chemotaxis machinery (T. Iino, Annu. Rev. Genet. 11:161-182, 1977; Y. Komeda, J. Bacteriol. 168: 1315-1318). The master operon, flbB, has been sequenced, and one of its gene products (FlaI) has been identified. On the basis of the deduced amino acid sequence, the FlbB protein has similarity to an alternate sigma factor which is responsible for expression of flagella in Bacillus subtilis. In addition, we have sequenced the 5' regions of a number of flagellar operons and compared these sequences with the 5' region of flagellar operons directly and indirectly under FlbB and FlaI control. We found both a consensus sequence which has been shown to be in all other flagellar operons (J. D. Helmann and M. J. Chamberlin, Proc. Natl. Acad. Sci. USA 84:6422-6424) and a derivative consensus sequence, which is found only in the 5' region of operons directly under FlbB and FlaI control.

Amino Acid Sequence↗

Nuclear proteins that bind the pre-mRNA 3' splice site sequence r(UUAG/G) and the human telomeric DNA sequence d(TTAGGG)n.

HeLa cell nuclear proteins that bind to single-stranded d(TTAGGG)n, the human telomeric DNA repeat, were identified and purified by a gel retardation assay. Immunological data and peptide sequencing experiments indicated that the purified proteins were identical or closely related to the heterogeneous nuclear ribonucleoproteins (hnRNPs) A1, A2-B1, D, and E and to nucleolin. These proteins bound to RNA oligonucleotides having r(UUAGGG) repeats more tightly than to DNA of the same sequence. The binding was sequence specific, as point mutation of any of the first 4 bases [r(UUAG)] abolished it. The fraction containing D and E hnRNPs was shown to bind specifically to a synthetic oligoribonucleotide having the 3' splice site sequence of the human beta-globin intervening sequence 1, which includes the sequence UUAGG. Proteins in this fraction were further identified by two-dimensional gel electrophoresis as D01, D02, D1*, and E0; intriguingly, these members of the hnRNP D and E groups are nuclear proteins that are not stably associated with hnRNP complexes. These studies establish the binding specificities of these D and E hnRNPs. Furthermore, they suggest the possibility that these hnRNPs could perhaps bind to chromosome telomeres, in addition to having a role in pre-mRNA metabolism.

Amino Acid Sequence↗