Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Comparative sequence and expression analyses of four mammalian VPS4 genes.

The VPS4 gene is a member of the AAA-family; it codes for an ATPase which is involved in lysosomal/endosomal membrane trafficking. VPS4 genes are present in virtually all eukaryotes. Exhaustive data mining of all available genomic databases from completely or partially sequenced organisms revealed the existence of up to three paralogues, VPS4a, -b, and -c. Whereas in the genome of lower eukaryotes like yeast only one VPS4 representative is present, we found that mammals harbour two paralogues, VPS4a and VPS4b. Most interestingly, the Fugu fish contains a third VPS4 paralogue (VPS4c). Sequence comparison of the three VPS4 paralogues indicates that the Fugu VPS4c displays sequence features intermediate between VPS4a and VPS4b. Using complete mammalian VPS4a and VPS4b cDNA clones as probes, genomic clones of both VPS4 paralogues in human and mouse were identified and sequenced. The chromosomal loci of all four VPS4 genes were determined by independent methods. A BLAST search of the human genome database with the human VPS4A sequence yielded a double match, most likely due to a faulty assembly of sequence contigs in the human draft sequence. Fluorescent in situ hybridization and radiation hybrid analyses demonstrated that human and mouse VPS4A/a and VPS4B/b are located on syntenic chromosomal regions. Northern blot and semi-quantitative reverse transcription analyses showed that mouse VPS4a and VPS4b are differentially expressed in different organs, suggesting that the two paralogues have developed different functional properties since their divergence. To investigate the subcellular distribution of the murine VPS4 paralogues, we transiently expressed various fluorescent VPS4 fusion proteins in mouse 3T3 cells. All tested VPS4 fusion proteins were found in the cytosol. Expression of dominant-negative mutant VPS4 fusion proteins led to their concentration in the perinuclear region. Co-expression of VPS4a-GFP and VPS4b-dsRed fusion proteins revealed a partial co-localization that was most prominent with mutant VPS4a and VPS4b proteins. A physical interaction between the mouse paralogues was also supported by two-hybrid analyses.

3T3 Cells↗

Array comparative genomic hybridization analysis of uterine leiomyosarcoma.

PURPOSE: Using a genome-wide array-based comparative genomic hybridization (array-CGH), DNA copy number changes in uterine leiomyosarcoma were analyzed. MATERIALS AND METHODS: We analyzed 4 cases of uterine leiomyoma and 7 cases of uterine leiomyosarcoma. The paraffin-fixed tissue samples were microdissected under microscope and DNA was extracted. Array-based CGH and fluorescence in situ hybridization (FISH) were carried out with Genome database (Gene Ontology). RESULTS: Uterine leiomyoma showed no genetic alterations, while all of 7 cases of uterine leiomyosarcoma showed specific gains and losses. The percentage of average gains and losses were 4.86% and 15.1%, respectively. The regions of high level of gain were 7q36.3, 7q33-q35, 12q13-12q15, and 12q23.3. And the regions of homozygous loss were 1p21.1, 2p22.2, 6p11.2, 9p21.1, 9p21.3, 9p22.1, 14q32.33, and 14q32.33 qter. There were no recurrent regions of gain, but recurrent regions of loss were 1p21.1-p21.2, 1p22.3-p31.1, 9p21.2-p22.2, 10q25-q25.2, 11q24.2-q25, 13q12-q12.13, 14q31.1-q31.3, 14q32.32-q32.33, 15q11-q12, 15q13-q14, 18q12.1-q12.2, 18q22.1-q22.3, 20p12.1, and 21q22.12-q22.13. In the high level of gain regions, BAC clones encoded HMGIC, SAS, MDM2, TIM1 genes. Frequently gained BAC clone-encoded genes were TIM1, PDGFR-beta, REC Q4, VAV2, FGF4, KLK2, PNUTL1, GDNF, FLG, EXT1, WISP1, HER-2, and SOX18. The genes encoded by frequently lost BAC clones were LEU1, ERCC5, THBS1, DCC, MBD2, SCCA1, FVT1, CYB5, and ETS2/E2. A subset of cellular processes from each gene was clustered by Gene Ontology database. CONCLUSION: Using array-CGH, chromosomal aberrations related to uterine leiomyosarcoma were identified. The high resolution of array-CGH combined with human genome database would give a chance to find out possible target genes present in the gained or lost clones.

Adult↗

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software↗

Identification and characterization of two parathyroid hormone-like molecules in zebrafish.

Zebrafish (Danio rerio) have receptors homologous to the human PTH (hPTH)/PTHrP receptor (PTH1R) and PTH-2 receptor (PTH2R) and an additional receptor (PTH3R) with high homology to the PTH1R. To find natural ligands for zPTH1R and zPTH3R, we searched the zebrafish genomic database and discovered two distinct regions that, when translated (zPTH1 and zPTH2), showed high homology to hPTH. Isolation of cDNAs and determination of the intron/exon boundaries revealed genomic structures which were similar to known PTHs. Peptides consisting of the first 34 amino acids after the pre- and prosequences of the zebrafish PTHs (zPTHs) were synthesized and were shown to be fully active at the hPTH1R. zPTH2(1-34) was, however, approximately 30-fold less potent at the zPTH1R than hPTH(1-34), hPTHrP(1-36), and zPTH1(1-34). When tested with zPTH3R, zPTH1(1-34) and hPTHrP(1-36) showed similar potencies, whereas the potency of zPTH2(1-34) was moderately (3-fold) reduced. To determine whether other fishes have multiple PTHs, we searched the genomic database of the Japanese pufferfish (Takifugu rubripes) and identified zPTH1 and zPTH2 homologs. Phylogenetic analysis showed that PTHs from zebrafish and pufferfish are more closely related to each other than to known mammalian PTH homologs or to PTHrP and tuberoinfundibular peptide of 39 residues. This is consistent with evolution of two teleost PTH-like peptides occurring after the evolutionary divergence between fishes and mammals. Overall, the PTH system appears more complex in fishes than in mammals, providing evidence of continued evolution in nontetrapod species. The availability of multiple forms of fish PTH and their receptors provide additional tools for PTH ligand/receptor structure-function studies.

Amino Acid Sequence↗

The plant structure ontology, a unified vocabulary of anatomy and morphology of a flowering plant.

Formal description of plant phenotypes and standardized annotation of gene expression and protein localization data require uniform terminology that accurately describes plant anatomy and morphology. This facilitates cross species comparative studies and quantitative comparison of phenotypes and expression patterns. A major drawback is variable terminology that is used to describe plant anatomy and morphology in publications and genomic databases for different species. The same terms are sometimes applied to different plant structures in different taxonomic groups. Conversely, similar structures are named by their species-specific terms. To address this problem, we created the Plant Structure Ontology (PSO), the first generic ontological representation of anatomy and morphology of a flowering plant. The PSO is intended for a broad plant research community, including bench scientists, curators in genomic databases, and bioinformaticians. The initial releases of the PSO integrated existing ontologies for Arabidopsis (Arabidopsis thaliana), maize (Zea mays), and rice (Oryza sativa); more recent versions of the ontology encompass terms relevant to Fabaceae, Solanaceae, additional cereal crops, and poplar (Populus spp.). Databases such as The Arabidopsis Information Resource, Nottingham Arabidopsis Stock Centre, Gramene, MaizeGDB, and SOL Genomics Network are using the PSO to describe expression patterns of genes and phenotypes of mutants and natural variants and are regularly contributing new annotations to the Plant Ontology database. The PSO is also used in specialized public databases, such as BRENDA, GENEVESTIGATOR, NASCArrays, and others. Over 10,000 gene annotations and phenotype descriptions from participating databases can be queried and retrieved using the Plant Ontology browser. The PSO, as well as contributed gene associations, can be obtained at www.plantontology.org.

Gene Expression Regulation, Plant↗

Searching for candidate genes for male infertility.

AIM: We describe an approach to search for candidate genes for male infertility using the two human genome databases: the public University of California at Santa Cruz (UCSC) and private Celera databases which list known and predicted gene sequences and provide related information such as gene function, tissue expression, known mutations and single nucleotide polymorphisms (SNPs). METHODS AND RESULTS: To demonstrate this in silico research, the following male infertility candidate genes were selected: (1) human BOULE, mutations of which may lead to germ cell arrest at the primary spermatocyte stage, (2) mutations of casein kinase 2 alpha genes which may cause globozoospermia, (3) DMR-N9 which is possibly involved in the spermatogenic defect of myotonic dystrophy and (4) several testes expressed genes at or near the breakpoints of a balanced translocation associated with hypospermatogenesis. We indicate how information derived from the human genome databases can be used to confirm these candidate genes may be pathogenic by studying RNA expression in tissue arrays using in situ hybridization and gene sequencing. CONCLUSION: The paper explains the new approach to discovering genetic causes of male infertility using information about the human genome.

Casein Kinase II↗

A complete survey of Trichoderma chitinases reveals three distinct subgroups of family 18 chitinases.

Genome-wide analysis of chitinase genes in the Hypocrea jecorina (anamorph: Trichoderma reesei) genome database revealed the presence of 18 ORFs encoding putative chitinases, all of them belonging to glycoside hydrolase family 18. Eleven of these encode yet undescribed chitinases. A systematic nomenclature for the H. jecorina chitinases is proposed, which designates the chitinases corresponding to their glycoside hydrolase family and numbers the isoenzymes according to their pI from Chi18-1 to Chi18-18. Phylogenetic analysis of H. jecorina chitinases, and those from other filamentous fungi, including hypothetical proteins of annotated fungal genome databases, showed that the fungal chitinases can be divided into three groups: groups A and B (corresponding to class V and III chitinases, respectively) also contained the so Trichoderma chitinases identified to date, whereas a novel group C comprises high molecular weight chitinases that have a domain structure similar to Kluyveromyces lactis killer toxins. Five chitinase genes, representing members of groups A-C, were cloned from the mycoparasitic species H. atroviridis (anamorph: T. atroviride). Transcription of chi18-10 (belonging to group C) and chi18-13 (belonging to a novel clade in group B) was triggered upon growth on Rhizoctonia solani cell walls, and during plate confrontation tests with the plant pathogen R. solani. Therefore, group C and the novel clade in group B may contain chitinases of potential relevance for the biocontrol properties of Trichoderma.

3' Untranslated Regions↗

Identification and disruption of the gene encoding the third member of the low-molecular-mass rhoptry complex in Plasmodium falciparum.

The low-molecular-mass rhoptry complex of Plasmodium falciparum consists of three proteins, rhoptry-associated protein 1 (RAP1), RAP2, and RAP3. The genes encoding RAP1 and RAP2 are known; however, the RAP3 gene has not been identified. In this study we identify the RAP3 gene from the P. falciparum genome database and show that this protein is part of the low-molecular-mass rhoptry complex. Disruption of RAP3 demonstrated that it is not essential for merozoite invasion, probably because RAP2 can complement the loss of RAP3. RAP3 has homology with RAP2, and the genes are encoded on chromosome 5 in a head-to-tail fashion. Analysis of the genome databases has identified homologous genes in all Plasmodium spp., suggesting that this protein plays a role in merozoite invasion. The region surrounding the RAP3 homologue in the Plasmodium yoelii genome is syntenic with the same region in P. falciparum; however, there is a single gene. Phylogenetic comparison of the RAP2/3 protein family from Plasmodium spp. suggests that the RAP2/3 duplication occurred after divergence of these parasite species.

Amino Acid Sequence↗

LTR_STRUC: a novel search and identification program for LTR retrotransposons.

MOTIVATION: Long terminal repeat (LTR) retrotransposons constitute a substantial fraction of most eukaryotic genomes and are believed to have a significant impact on genome structure and function. Conventional methods used to search for LTR retrotransposons in genome databases are labor intensive. We present an efficient, reliable and automated method to identify and analyze members of this important class of transposable elements. RESULTS: We have developed a new data-mining program, LTR_STRUC (LTR retrotransposon structure program) which identifies and automatically analyzes LTR retrotransposons in genome databases by searching for structural features characteristic of such elements. LTR_STRUC has significant advantages over conventional search methods in the case of LTR retrotransposon families having low sequence homology to known queries or families with atypical structure (e.g. non-autonomous elements lacking canonical retroviral ORFs) and is thus a discovery tool that complements established methods. LTR_STRUC finds LTR retrotransposons using an algorithm that encompasses a number of tasks that would otherwise have to be initiated individually by the user. For each LTR retrotransposon found, LTR_STRUC automatically generates an analysis of a variety of structural features of biological interest. AVAILABILITY: The LTR_STRUC program is currently available as a console application free of charge to academic users from the authors.

Algorithms↗

Involvement of some large immunophilins and their ligands in the protection and regeneration of neurons: a hypothetical mode of action.

The powerful immunosuppressive drugs such as FK506 and its derivatives induce some regeneration and protection of neurons from ischaemic brain injury and some other neurological disorders. The drugs form complexes with diverse FKBPs but apparently the FKBP52/FK506 complex was shown to be involved in the protection and regeneration of neurons. We used several different sequence attributes in searching diverse genomic databases for similar motifs as those present in the FKBPs. A Fortran library of algorithms (Par_Seq) has been designed and used in searching for the similarity of sequence motifs extracted from the multiple sequence alignments of diverse groups of proteins (query motifs) and the target motifs which are encoded in various genomes. The following sequence attributes were used in the establishment of the degree of convergence between: (A) amino acid (AA) sequence similarity (ID) of the query/target motifs and (B) their: (1) AA composition (AAC); (2) hydrophobicity (HI); (3) Jensen-Shannon entropy; and (4) AA propensity to form a particular secondary structure. The sequence hallmark of two different groups of peptidylprolyl cis/trans isomerases (PPIases), namely tetratricopetide repeat (TPR) motifs, which are present in the heat-shock cyclophilins and in the large FK506-binding proteins (FKBPs) were used to search various genomic databases. The Par_Seq algorithm has revealed that the TPR motifs have similar sequence attributes as a number of hydrophobic sequence segments of functionally unrelated membrane proteins, including some of the TMs from diverse G protein-coupled receptors (GPCRs). It is proposed that binding of the FKBP52/FK506 complex to the membranes via the TPR motifs and its interaction with some membrane proteins could be in part responsible for some neuro-regeneration and neuro-protection of the brain during some ischaemia-induced stresses.

Algorithms↗

Comparative proteomics of the Mycobacterium leprae binding protein myelin P0: its implication in leprosy and other neurodegenerative diseases.

Mycobacterium leprae, the causative agent of leprosy invades Schwann cells of the peripheral nerves leading to nerve damage and disfigurement, which is the hallmark of the disease. Wet experiments have shown that M. leprae binds to a major peripheral nerve protein, the myelin P zero (P0). This protein is specific to peripheral nerve and may be important in the initial step of M. leprae binding and invasion of Schwann cells which is the feature of leprosy. Though the receptors on Schawann cells, cytokines, chemokines and antibodies to M. leprae have been identified the molecular mechanism of nerve damage and neurodegeneration is not clearly defined. Recently pathogen and host protein/nucleotide sequence similarities (molecular mimicry) have been implicated in neurodegenerative diseases. The approach of the present study is to utilise bioinformatic tools to understand leprosy nerve damage by carrying out sequence and structural similarity searches of myelin P0 with leproma and other genomic database. Since myelin P0 is unique to peripheral nerve, its sequence and structural similarities in other neuropathogens have also been noted. Comparison of myelin P0 with the M. leprae proteins revealed two characterised proteins, Ferrodoxin NADP reductase and a conserved membrane protein, which showed similarity to the query sequence. Comparison with the entire genomic database (www.ncbi.nlm.nih.gov) by basic local alignment search tool for proteins (BLASTP) and fold classification of structure-structure alignment of proteins (FSSP) searches revealed that myelin P0 had sequence/structural similarities to the poliovirus receptor, coxsackie-adenovirus receptor, anthrax protective antigen, diphtheria toxin, herpes simplex virus, HIV gag-1 peptide, and gp120 among others. These proteins are known to be associated directly or indirectly with neruodegeneration. Sequence and structural similarities to the immunoglobin regions of myelin P0 could have implications in host-pathogen interactions, as it has homophilic adhesive properties. Although these observed similarities are not highly significant in their percentage identity, they could be functionally important in molecular mimicry, receptor binding and cell signaling events involved in neurodegeneration.

Amino Acid Sequence↗

Identification & characterisation of the two novel streptococcal pyrogenic exotoxins SPE-L & SPE-M.

BACKGROUND & OBJECTIVES: The streptococcal pyrogenic exotoxins (SPEs) are produced by Streptococcus pyogenes and belong to the family of bacterial superantigens, a group of highly mitogenic proteins. The aim of this study was to search unfinished streptococcal genomes for novel superantigens, to generate recombinant proteins from potential open reading frames (ORFs) and to analyse them for superantigen activity. METHODS: The microbial genome database was searched using a TBLASTN search programme. Genotyping of S. equi and S. pyogenes isolates was done using the specific primer pairs. The spe-l and spe-m genes were amplified by PCR. RESULTS: Two novel streptococcal superantigen genes (sepe-l and sepe-m) were identified from the Streptococcus equi genomic database at the Sanger Centre. Genotyping of S. pyogenes isolates resulted in the detection of the orthologous genes spe-l and spe-m in a restricted number of S. pyogenes isolates and revealed a link of spe-l to the M89 serotype. Recombinant SPE-L and rSPE-M were highly mitogenic for human peripheral blood lymphocytes with half maximum responses at 1 pg/ml and 10 pg/ml, respectively. The results from competitive binding experiments suggest that both proteins bind MHC class II at the beta-chain, but not at the alpha-chain. The most common target for both toxins were human Vbetal.1 expressing T cells. Seroconversion against SPE-L and SPE-M was observed in healthy blood donors. INTERPRETATION & CONCLUSION: The two novel ORFs identified in both, S. equi and S. pyogenes, code for proteins that show typical superantigen features. The seroconversion seen in some blood donors suggest that the proteins are indeed produced by the bacteria. Interestingly, the spe-l gene is highly associated with S. pyogenes M89, which is linked to acute rheumatic fever in New Zealand.

Bacterial Proteins↗

Diversity and biocatalytic potential of epoxide hydrolases identified by genome analysis.

Epoxide hydrolases play an important role in the biodegradation of organic compounds and are potentially useful in enantioselective biocatalysis. An analysis of various genomic databases revealed that about 20% of sequenced organisms contain one or more putative epoxide hydrolase genes. They were found in all domains of life, and many fungi and actinobacteria contain several putative epoxide hydrolase-encoding genes. Multiple sequence alignments of epoxide hydrolases with other known and putative alpha/beta-hydrolase fold enzymes that possess a nucleophilic aspartate revealed that these enzymes can be classified into eight phylogenetic groups that all contain putative epoxide hydrolases. To determine their catalytic activities, 10 putative bacterial epoxide hydrolase genes and 2 known bacterial epoxide hydrolase genes were cloned and overexpressed in Escherichia coli. The production of active enzyme was strongly improved by fusion to the maltose binding protein (MalE), which prevented inclusion body formation and facilitated protein purification. Eight of the 12 fusion proteins were active toward one or more of the 21 epoxides that were tested, and they converted both terminal and nonterminal epoxides. Four of the new epoxide hydrolases showed an uncommon enantiopreference for meso-epoxides and/or terminal aromatic epoxides, which made them suitable for the production of enantiopure (S,S)-diols and (R)-epoxides. The results show that the expression of epoxide hydrolase genes that are detected by analyses of genomic databases is a useful strategy for obtaining new biocatalysts.

Animals↗

Molecular cloning, genomic organization, chromosomal mapping and subcellular localization of mouse PAP7: a PBR and PKA-RIalpha associated protein.

A mouse protein that interacts with the peripheral-type benzodiazepine receptor (PBR) and the cAMP-dependent protein kinase A (PKA) regulatory subunit RIalpha (PKA-RIalpha), named PBR and PKA associated protein 7 (PAP7) was identified and shown to be involved in hormone-induced steroid biosynthesis in testicular Leydig cells. In the present study, mouse PAP7 cDNA was extended by 5'-rapid amplification of cDNA ends; and a 3432 bp sequence, encoding a 525-amino-acid protein with a calculated molecular weight of 60 kDa, was re-assembled. Mouse and human PAP7 share an 85% amino acid identity and contain a conserved acyl-CoA-binding protein/diazepam binding inhibitor (ACBP/DBI) motif. ACBP/DBI has been identified as the endogenous PBR ligand able to stimulate mitochondrial steroid formation in all steroidogenic cells. The full-length mouse PAP7 gene was cloned and assembled by screening a BAC clone, polymerase chain reaction and searching the mouse genome database. The gene is approximately 29 kb in length and includes eight exons and seven introns. Although it is shorter than the human PAP7 gene, all exons are conserved between the mouse and human. The mouse PAP7 gene was mapped to chromosome 1H3-5 by fluorescence in situ hybridization in agreement with in silico search of the mouse genome database that mapped the PAP7 cDNA sequence to the 1H4 area. Immunofluorescence confocal microscopy demonstrated that PAP7 is mainly localized in the trans-Golgi apparatus and mitochondria in mouse tumor Leydig cells, in agreement with its proposed function in targeting the PKA isoenzyme to organelles rich in PBR, i.e. mitochondria, where phosphorylation of specific protein substrates mediates the hormone-induced steroid formation.

Adaptor Proteins, Signal Transducing↗

Exploiting conserved structure for faster annotation of non-coding RNAs without loss of accuracy.

MOTIVATION: Non-coding RNAs (ncRNAs)-functional RNA molecules not coding for proteins-are grouped into hundreds of families of homologs. To find new members of an ncRNA gene family in a large genome database, covariance models (CMs) are a useful statistical tool, as they use both sequence and RNA secondary structure information. Unfortunately, CM searches are slow. Previously, we introduced 'rigorous filters', which provably sacrifice none of CMs' accuracy, although often scanning much faster. A rigorous filter, using a profile hidden Markov model (HMM), is built based on the CM, and filters the genome database, eliminating sequences that provably could not be annotated as homologs. The CM is run only on the remainder. Some biologically important ncRNA families could not be scanned efficiently with this technique, largely due to the significance of conserved secondary structure relative to primary sequence in identifying these families. Current heuristic filters are also expected to perform poorly on such families. RESULTS: By augmenting profile HMMs with limited secondary structure information, we obtain rigorous filters that accelerate CM searches for virtually all known ncRNA families from the Rfam Database and tRNA models in tRNAscan-SE. These filters scan an 8 gigabase database in weeks instead of years, and uncover homologs missed by heuristic techniques to speed CM searches. AVAILABILITY: Software in development; contact the authors.

Algorithms↗

A Saccharomyces cerevisiae Internet protein resource now available.

The QUEST Protein Database Center is now making available two Saccharomyces cerevisiae protein databases via the Internet. The yeast electrophoretic protein database (YEPD) is a database of approximately one hundred protein identifications on two-dimensional gels. The yeast protein database (YPD) is a database of gene names and properties of over 3500 yeast proteins of known sequence. These databases can be accessed via a World-Wide Web (WWW) server (URL http:@siva.cshl.org). YPD is available via public ftp (isis.cshl.org) as well, in a spreadsheet format, and in ASCII format. When accessed via WWW, both of these databases have hypertext links to other biological data, such as the SWISS-PROT protein sequence database and the Saccharomyces Genome Database (SacchDB), and to each other.

Computer Communication Networks↗

Identification of new eukaryotic tRNA genes in genomic DNA databases by a multistep weight matrix analysis of transcriptional control regions.

A linear method for the search of eukaryotic nuclear tRNA genes in DNA databases is described. Based on a modified version of the general weight matrix procedure, our algorithm relies on the recognition of two intragenic control regions known as A and B boxes, a transcription termination signal, and on the evaluation of the spacing between these elements. The scanning of the eukaryotic nuclear DNA database using this search algorithm correctly identified 933 of the 940 known tRNA genes (0.74% of false negatives). Thirty new potential tRNA genes were identified, and the transcriptional activity of two of them was directly verified by in vitro transcription. The total false positive rate of the algorithm was 0.014%. Structurally unusual tRNA genes, like those coding for selenocysteine tRNAs, could also be recognized using a set of rules concerning their specific properties, and one human gene coding for such tRNA was identified. Some of the newly identified tRNA genes were found in rather uncommon genomic positions: 2 in centromeric regions and 3 within introns. Furthermore, the presence of extragenically located B boxes in tRNA genes from various organisms could be detected through a specific subroutine of the standard search program.

Algorithms↗

Structure analysis of two Toxoplasma gondii and Neospora caninum satellite DNA families and evolution of their common monomeric sequence.

A family of repetitive DNA elements of approximately 350 bp-Sat350-that are members of Toxoplasma gondii satellite DNA was further analyzed. Sequence analysis identified at least three distinct repeat types within this family, called types A, B, and C. B repeats were divided into the subtypes B1 and B2. A search for internal repetitions within this family permitted the identification of conserved regions and the design of PCR primers that amplify almost all these repetitive elements. These primers amplified the expected 350-bp repeats and a novel 680-bp repetitive element (Sat680) related to this family. Two additional tandemly repeated high-order structures corresponding to this satellite DNA family were found by searching the Toxoplasma genome database with these sequences. These studies were confirmed by sequence analysis and identified: (1). an arrangement of AB1CB2 350-bp repeats and (2). an arrangement of two 350-bp-like repeats, resulting in a 680-bp monomer. Sequence comparison and phylogenetic analysis indicated that both high-order structures may have originated from the same ancestral 350-bp repeat. PCR amplification, sequence analysis and Southern blot showed that similar high-order structures were also found in the Toxoplasma-sister taxon Neospora caninum. The Toxoplasma genome database (http://ToxoDB.org ) permitted the assembly of a contig harboring Sat350 elements at one end and a long nonrepetitive DNA sequence flanking this satellite DNA. The region bordering the Sat350 repeats contained two differentially expressed sequence-related regions and interstitial telomeric sequences.

Animals↗