Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Ribosomal ITS2 sequence data for Anopheles maculipennis and An. messeae in northern Greece, with a critical assessment of previously published sequences.

DNA sequences were generated for eight specimens of the Anopheles maculipennis complex from Florina in NW Greece, and identified to species on the basis of comparison with ITS2 sequences for members of the complex already in GenBank. The sequences revealed the presence of An. maculipennis and An. messeae in Florina. Problems with sequence reliability and accessibility of sequences generated in earlier studies of Palaearctic members of the complex are discussed.

Animals↗

Automatic generation of primary sequence patterns from sets of related protein sequences.

We have developed a computer algorithm that can extract the pattern of conserved primary sequence elements common to all members of a homologous protein family. The method involves clustering the pairwise similarity scores among a set of related sequences to generate a binary dendrogram (tree). The tree is then reduced in a stepwise manner by progressively replacing the node connecting the two most similar termini by one common pattern until only a single common "root" pattern remains. A pattern is generated at a node by (i) performing a local optimal alignment on the sequence/pattern pair connected by the node with the use of an extended dynamic programming algorithm and then (ii) constructing a single common pattern from this alignment with a nested hierarchy of amino acid classes to identify the minimal inclusive amino acid class covering each paired set of elements in the alignment. Gaps within an alignment are created and/or extended using a "pay once" gap penalty rule, and gapped positions are converted into gap characters that function as 0 or 1 amino acid of any type during subsequent alignment. This method has been used to generate a library of covering patterns for homologous families in the National Biomedical Research Foundation/Protein Identification Resource protein sequence data base. We show that a covering pattern can be more diagnostic for sequence family membership than any of the individual sequences used to construct the pattern.

Amino Acid Sequence↗

Nucleotide sequence of a cDNA clone (SUBT1) partly homologous to a human subtelomeric repeat sequence.

Telomeres contains short tandem repeats (TTAGGG)n that are usually several kilobases in size. DNA sequences located next to the repeats are called telomere-associated or subtelomeric sequences. In this report we describe isolation and characterization of a new cDNA sequence isolated with a telomere specific, chromosome 17 genomic clone. No open reading frame of notable length could be detected. However comparison with sequences in the EMBL/Genbank revealed a 897 bp of close to perfect match, in the 3'end with the centromeric end of a genomic clone containing subtelomeric repeat sequences. The potential function of the SUBT1 transcript remains to be elucidated.

Chromosomes, Human, Pair 17↗

Cloning and sequencing of the rat preproepidermal growth factor cDNA: comparison with mouse and human sequences.

We have cloned and sequenced the cDNA corresponding to the rat preproepidermal growth factor (ppEGF) mRNA. The cDNA contained 4,801 nucleotides, similar to that reported for the mouse (4,749 nucleotides) and the human mRNAs (4,871 nucleotides). The predicted protein sequence would contain 1,133 amino acids, smaller than that reported for the mouse (1,217 amino acids) and the human sequences (1,207 amino acids). The results of the sequencing of several cDNA clones suggested the existence of more than one structural gene for ppEGF. In addition, there was an occurrence of alternative splicing events, resulting in deletions of entire exons from the mature mRNA. These alternative splicing events do not create frameshift mutations but cause a deletion of one or more of the "EGF-like" repeat units from the ppEGF. There is approximately the same homology between the rat and mouse amino acid sequences both in the EGF region and in the other regions of the ppEGF protein. We conclude that, because of this conservation of homology, there may be an important function performed by these other regions of the ppEGF besides their function as a precursor for the EGF protein.

Amino Acid Sequence↗

Sequence complexity profiles of prokaryotic genomic sequences: a fast algorithm for calculating linguistic complexity.

MOTIVATION: One of the major features of genomic DNA sequences, distinguishing them from texts in most spoken or artificial languages, is their high repetitiveness. Variation in the repetitiveness of genomic texts reflects the presence and density of different biologically important messages. Thus, deviation from an expected number of repeats in both directions indicates a possible presence of a biological signal. Linguistic complexity corresponds to repetitiveness of a genomic text, and potential regulatory sites may be discovered through construction of typical patterns of complexity distribution. RESULTS: We developed software for fast calculation of linguistic sequence complexity of DNA sequences. Our program utilizes suffix trees to compute the number of subwords present in genomic sequences, thereby allowing calculation of linguistic complexity in time linear in genome size. The measure of linguistic complexity was applied to the complete genome of Haemophilus influenzae. Maps of complexity along the entire genome were obtained using sliding windows of 40, 100, and 2000 nucleotides. This approach provided an efficient way to detect simple sequence repeats in this genome. In addition, local profiles of complexity distribution around the starts of translation were constructed for 21 complete prokaryotic genomes. We hypothesize that complexity profiles correspond to evolutionary relationships between organisms. We found principal differences in profiles of the GC-rich and other (non-GC-rich) genomes. We also found characteristic differences in profiles of AT genomes, which probably reflect individual species variations in translational regulation. AVAILABILITY: The program is available upon request from Alexander Bolshoy or at http://csweb.haifa.ac.il/library/#complex.

Algorithms↗

A software program combining sequence motif searches with keywords for finding repeats containing DNA sequences.

MOTIVATION: One of the most interesting features of genomes (both coding and non-coding regions) is the presence of relatively short tandemly repeated DNA sequences known as tandem repeats (TRs). We developed a new PC-based stand-alone software analysis program, combining sequence motif searches with keywords such as organs, tissues, cell lines or development stages for finding exact, inexact and compound, TRs. Tandem Repeats Analyzer 1.5 (TRA) has several advanced repeat search parameters/options over other repeat finder programs as it does not only accept GenBank, FASTA and expressed sequence tag (EST) sequence files but also does analysis of multifiles with multisequences. Advanced user-defined parameters/options let the researchers use different motif lengths search criteria for varying motif lengths simultaneously. The outputs show statistical results to be evaluated by the user. The discovery of TRs in ESTs could be useful for both gene mapping and association studies and discovering TRs located in coding regions of important genes that are expressed under various conditions of environment, stress, organ, tissue and development stage. RESULTS: In this paper, we demonstrated applications of TRA using 175 899 ESTs sequences for three Arabidopsis spp. downloaded from GenBank. The EST-SSRs/ESTs ratios were found 43.1%, 15.3% and 2.34% in A.lyrata, A.thaliana and A.halleri, respectively. Analysis revealed that organs, tissues and development stages possessed different amounts of repeats and repeat compositions. This indicated that the distribution of TRs among the tissues or organs may not be random differing from the untranscribed repeats found in genomes. AVAILABILITY: The program can be obtained free by anonymous FTP from ftp.akdeniz.edu.tr/Araclar/TRA.

Arabidopsis↗

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.

MOTIVATION: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment. RESULTS: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity <30%). Its accuracy for aligning homologs (sequence identity >30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0. AVAILABILITY: The SPEM server and its executables are available on http://theory.med.buffalo.edu.

Algorithms↗

DynaPred: a structure and sequence based method for the prediction of MHC class I binding peptide sequences and conformations.

MOTIVATION: The binding of endogenous antigenic peptides to MHC class I molecules is an important step during the immunologic response of a host against a pathogen. Thus, various sequence- and structure-based prediction methods have been proposed for this purpose. The sequence-based methods are computationally efficient, but are hampered by the need of sufficient experimental data and do not provide a structural interpretation of their results. The structural methods are data-independent, but are quite time-consuming and thus not suited for screening of whole genomes. Here, we present a new method, which performs sequence-based prediction by incorporating information obtained from molecular modeling. This allows us to perform large databases screening and to provide structural information of the results. RESULTS: We developed a SVM-trained, quantitative matrix-based method for the prediction of MHC class I binding peptides, in which the features of the scoring matrix are energy terms retrieved from molecular dynamics simulations. At the same time we used the equilibrated structures obtained from the same simulations in a simple and efficient docking procedure. Our method consists of two steps: First, we predict potential binders from sequence data alone and second, we construct protein-peptide complexes for the predicted binders. So far, we tested our approach on the HLA-A0201 allele. We constructed two prediction models, using local, position-dependent (DynaPred(POS)) and global, position-independent (DynaPred) features. The former model outperformed the two sequence-based methods used in our evaluation; the latter shows a much higher generalizability towards other alleles than the position-dependent models. The constructed peptide structures can be refined within seconds to structures with an average backbone RMSD of 1.53 A from the corresponding experimental structures.

Algorithms↗

Complete chloroplast genome sequences from Korean ginseng (Panax schinseng Nees) and comparative analysis of sequence evolution among 17 vascular plants.

The nucleotide sequence of Korean ginseng (Panax schinseng Nees) chloroplast genome has been completed (AY582139). The circular double-stranded DNA, which consists of 156,318 bp, contains a pair of inverted repeat regions (IRa and IRb) with 26,071 bp each, which are separated by small and large single copy regions of 86,106 bp and 18,070 bp, respectively. The inverted repeat region is further extended into a large single copy region which includes the 5' parts of the rpsl9 gene. Four short inversions associated with short palindromic sequences that form stem-loop structures were also observed in the chloroplast genome of P. schinseng compared to that of Nicotiana tabacum. The genome content and the relative positions of 114 genes (75 peptide-encoding genes, 30 tRNA genes, 4 rRNA genes, and 5 conserved open reading frames [ycfs]), however, are identical with the chloroplast DNA of N. tabacum. Sixteen genes contain one intron while two genes have two introns. Of these introns, only one (trnL-UAA) belongs to the self-splicing group I; all remaining introns have the characteristics of six domains belonging to group II. Eighteen simple sequence repeats have been identified from the chloroplast genome of Korean ginseng. Several of these SSR loci show infra-specific variations. A detailed comparison of 17 known completed chloroplast genomes from the vascular plants allowed the identification of evolutionary modes of coding segments and intron sequences, as well as the evaluation of the phylogenetic utilities of chloroplast genes. Furthermore, through the detailed comparisons of several chloroplast genomes, evolutionary hotspots predominated by the inversion end points, indel mutation events, and high frequencies of base substitutions were identified. Large-sized indels were often associated with direct repeats at the end of the sequences facilitating intra-molecular recombination.

Base Sequence↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. I. Sequence features in the 1 Mb region from map positions 64% to 92% of the genome.

The contiguous sequence of 1,003,450 bp spanning map positions 64% to 92% of the genome of Synechocystis sp. strain PCC6803 has been deduced. Computer analysis of the sequence predicts that this region contains at least 818 potential ORFs, in which 255 (31%) were either genes that had already been identified or their homologues, 84 (10%) were homologues to registered hypothetical genes, and 149 (18%) showed weak similarities to reported genes. The remaining 330 ORFs showed no apparent similarity to any reported genes or carried no significant protein motifs. The potential ORFs as a whole occupied 86% of the sequenced region, implying compact arrangement of genes in the genome. As to the structural RNA genes, one rRNA operon consisting of 5,028 bp and at least 11 species of tRNA genes were identified. It is noteworthy that 10 out of the 11 tRNA species showed significant sequence similarities to tRNAs reported in plant chloroplasts. As other notable unique sequences, three classes of IS-like elements each with characteristics typical of IS elements were identified, and a typical unit of WD(Trp-Asp)-repeats which have only been detected in the regulatory proteins of eukaryotes was identified within the large 5,079-bp ORF located at map position 69%.

Base Sequence↗

Prediction of the coding sequences of unidentified human genes. V. The coding sequences of 40 new genes (KIAA0161-KIAA0200) deduced by analysis of cDNA clones from human cell line KG-1.

As part of our continuing efforts to accumulate information on the coding region of unidentified human genes, we newly determined the sequences of 40 cDNA clones of human cell line KG-1 which correspond to relatively long and nearly full-length transcripts, and predicted the coding sequences of the corresponding genes, named KIAA0161 to 0200. The average size of the cDNA clones analyzed was approximately 5.0 kb. A computer search of the sequences in public databases indicated that the sequences of 20 genes were unrelated to any reported genes, while the remaining 20 genes carried sequences which show some similarities to known genes. Among the genes in the latter category, KIAA0167 contained a Zn-finger motif with significant structural similarity to that of the yeast transcription factor GCS1, and KIAA0189 was classified into the RhoGAP gene family. Stretches of typical CAG (Gln) repeats, which were often correlated with genetic disorders, were found in KIAA0181 and KIAA0192. Another novel repeat composed of alternating Arg and Glu was identified in KIAA0182. Northern hybridization analysis demonstrated that 10 genes are expressed in a cell- or tissue-specific manner.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. XVIII. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

In our series of human cDNA projects for accumulating sequence information on the coding sequences of unidentified genes, we herein present the entire sequences of 100 cDNA clones of unidentified genes, named KIAA1544 to KIAA1643, from two sets of size-fractionated human adult and fetal brain cDNA libraries. The average sizes of the inserts and corresponding open reading frames of cDNA clones analyzed here reached 4.6 kb and 2.8 kb (930 amino acid residues), respectively. By computer-assisted database search of the deduced amino acid sequences, 48 predicted gene products were classified into the five functional categories of proteins relating to cell signaling/communication, nucleic acid management, cell structure/motility, protein management and metabolism. Homology search against the databases for proteins deduced from yeast, nematode and fly full genome sequences revealed only one gene (KIAA1630) was entirely conserved among human and these three organisms in the 100 genes reported here. Additionally, their chromosomal loci were determined by using human-rodent hybrid panels unless they were already assigned in the public databases. Furthermore, the expression profiles of the genes were also studied in 10 human tissues, 8 brain regions, spinal cord, fetal brain and fetal liver by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Amino Acid Sequence↗

Prediction of the coding sequences of unidentified human genes. XX. The complete sequences of 100 new cDNA clones from brain which code for large proteins in vitro.

To accumulate information on the coding sequences of unidentified genes, we have carried out a sequencing project of human cDNA clones which encode large proteins. We herein present the entire sequences of 100 cDNA clones of unidentified human genes, named KIAA1776 and KIAA1780-KIAA1878, from size-fractionated cDNA libraries derived from human fetal brain, adult whole brain, hippocampus and amygdala. Most of the cDNA clones to be entirely sequenced were selected as cDNAs which were shown to have coding potentiality by in vitro transcription/translation experiments, and some clones were chosen by using computer-assisted analysis of terminal sequences of cDNAs. Three of these clones (fibrillin2/KIAA1776, MEGF10/KIAA1780 and MEGF11/KIAA1781) were isolated as genes encoding proteins with multiple EGF-like domains by motif-trap screening. The average sizes of the inserts and corresponding open reading frames of eDNA clones analyzed here reached 4.7 kb and 2.4 kb (785 amino acid residues), respectively. From the results of homology and motif searches against the public databases, the functional categories of the predicted gene products of 54 genes were determined; 93% of these predicted gene products (50 gene products) were classified as proteins related to cell signaling/communication, nucleic acid management, or cell structure/motility. To collect additional information on these genes, their expression profiles were also studied in 10 human tissues, 8 brain regions, spinal cord, fetal brain and fetal liver by reverse transcription-coupled polymerase chain reaction, products of which were quantified by enzyme-linked immunosorbent assay.

Adult↗

Characterization of minisatellites in Arabidopsis thaliana with sequence similarity to the human minisatellite core sequence.

A strategy based on random PCR amplification was used to isolate new repetitive elements of Arabidopsis thaliana. One of the random PCR product analyzed by this approach contained a tandem repetitive minisatellite sequence composed of 33 bp repeated units. The genomic locus corresponding to this PCR product was isolated by screening a lambda genomic library. New related loci were also isolated from the genomic library by screening with a 14 mer oligonucleotide representing a region conserved among the different repeated units. Alignment of the consensus sequence for each minisatellite locus allowed the definition of an Arabidopsis thaliana core sequence that shows strong sequence similarities with the human core sequence and with the generalized recombination signal Chi of Escherichia coli. The minisatellites were tested for their ability to detect polymorphism, and their chromosomal position was established.

Arabidopsis↗

The tandem repeat AGGGTAGGGT is, in the fission yeast, a proximal activation sequence and activates basal transcription mediated by the sequence TGTGACTG.

Ribosomal protein (rp) genes in the fission yeast Schizosaccharomyces pombe display two highly conserved sequence elements in the promoter region. The molecular dissection of these promoters revealed that basal transcription is not based on a TATA element. The sequence which promotes basal transcription is the conserved sequence CAGTCACA or the inverted form TGTGACTG, called the homol D box. Upstream of the homol D box a tandem repeat AGGGTAGGGT or the inverted form ACCCTACCCT appears in some promoters, called homol E. This element functions in the proximal arrangement with homol D as an activation sequence. A compilation of homol D and homol E sequences identified in other S.pombe promoters revealed that several putative polymerase II and polymerase III promoters display a homol D box or the homol E/homol D arrangement.

Base Sequence↗

INFOGENE: a database of known gene structures and predicted genes and proteins in sequences of genome sequencing projects.

INFOGENE is a database of known and predicted gene structures with descriptions of basic functional signals and gene components. It provides a possibility to create compilations of sequences with a given gene feature as well as to accumulate and analyze predicted genes in finished and unfinished sequences from genome sequencing projects. Protein sequence similarity searches in the database of predicted proteins is offered through the BLASTP program. INFOGENE is realized under the Sequence Retrieval System that provides useful links with the other informational databases. The database is available through the WWW server of the Computational Genomics Group at http://genomic.sanger.ac.uk/db.html

Animals↗

Analysis of human immunodeficiency virus type 1 env and gag sequence variants derived from a mother and two vertically infected children provides evidence for the transmission of multiple sequence variants.

In order to investigate the transmission of human immunodeficiency virus type 1 (HIV-1) from mother-to-child we have examined serial plasma RNA samples obtained from a mother over an eight year period spanning four pregnancies. Child 1 and 2 (born January 1987 and June 1990) were uninfected whilst child 3 and 4 (born July 1992 and February 1994) were HIV positive. Genetic variation was examined within the viral population of the mother and her two infected children for both the V3 loop and flanking regions of the env gene and the p17 region of the gag gene. In one child (child 4) a highly homogeneous virus population was observed within both env and gag in contrast to the more heterogeneous virus population observed within the mother. Viral sequences of child 4 clustered within a single branch within the reconstructed phylogenetic tree. This is consistent with the transmission of a single maternal variant to the child in this case, which may indicate a selective process. By contrast, child 3 showed substantial genetic heterogeneity even within the first samples obtained shortly after birth. Sequences of child 3 clustered in two distinct groups within the phylogenetic tree and were separated by sequences of the mother. These results are not consistent with the selective transmission of a single maternal variant to the child in this case and we therefore propose that the infection within child 3 is the result of the transmission of multiple sequence variants to the child. All transmitted sequence variants were predicted to be of the macrophage-tropic, nonsyncytium-inducing (NSI) phenotype.

Amino Acid Sequence↗

The nucleotide sequence of the LPD1 gene encoding lipoamide dehydrogenase in Saccharomyces cerevisiae: comparison between eukaryotic and prokaryotic sequences for related enzymes and identification of potential upstream control sites.

The complete nucleotide sequence of the LPD1 gene, which encodes the lipoamide dehydrogenase component (E3) of the pyruvate dehydrogenase and 2-oxoglutarate dehydrogenase multienzyme complexes of Saccharomyces cerevisiae, has been established. The flanking region 5' to the LPD1 gene contains DNA sequences which show homology to known control sites found upstream of other yeast genes. The primary structure of the protein, determined from the DNA sequence, shows strong homology to a group of flavoproteins including Escherichia coli lipoamide dehydrogenase and pig heart lipoamide dehydrogenase. The amino acid sequence also reveals the presence of a potential targeting sequence at its N-terminus which may facilitate transport to and entry into mitochondria.

Amino Acid Sequence↗