Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A beta-1,3-N-acetylglucosaminyltransferase with poly-N-acetyllactosamine synthase activity is structurally related to beta-1,3-galactosyltransferases.

Human and mouse cDNAs encoding a new beta-1, 3-N-acetylglucosaminyltransferase (beta3GnT) have been isolated from fetal and newborn brain libraries. The human and mouse cDNAs included ORFs coding for predicted type II transmembrane polypeptides of 329 and 325 aa, respectively. The human and mouse beta3GnT homologues shared 90% similarity. The beta3GnT gene was widely expressed in human and mouse tissues, although differences in the transcript levels were visible, thus indicating possible tissue-specific regulation mechanisms. The beta3GnT enzyme showed a marked preference for Gal(beta1-4)Glc(NAc)-based acceptors, whereas no activity was detected on type 1 Gal(beta1-3)GlcNAc and O-glycan core 1 Gal(beta1-3)GalNAc acceptors. The new beta3GnT enzyme was capable of both initiating and elongating poly-N-acetyllactosamine chains, which demonstrated its identity with the poly-N-acetyllactosamine synthase enzyme (E.C. 2.4.1.149), showed no similarity with the i antigen beta3GnT enzyme described recently, and, strikingly, included several amino acid motifs in its protein that have been recently identified in beta-1,3-galactosyltransferase enzymes. The comparison between the new UDP-GlcNAc:betaGal beta3GnT and the three UDP-Gal:betaGlcNAc beta-1,3-galactosyltransferases-I, -II, and -III reveals glycosyltransferases that share conserved sequence motifs though exhibiting inverted donor and acceptor specificities. This suggests that the conserved amino acid motifs likely represent residues required for the catalysis of the glycosidic (beta1-3) linkage.

Amino Acid Sequence↗

Detecting patterns of protein distribution and gene expression in silico.

Most biological information is contained within gene and genome sequences. However, current methods for analyzing these data are limited primarily to the prediction of coding regions and identification of sequence similarities. We have developed a computer algorithm, CoSMoS (for context sensitive motif searches), which adds context sensitivity to sequence motif searches. CoSMoS was challenged to identify genes encoding peroxisome-associated and oleate-induced genes in the yeast Saccharomyces cerevisiae. Specifically, we searched for genes capable of encoding proteins with a type 1 or type 2 peroxisomal targeting signal and for genes containing the oleate-response element, a cis-acting element common to fatty acid-regulated genes. CoSMoS successfully identified 7 of 8 known PTS-containing peroxisomal proteins and 13 of 14 known oleate-regulated genes. More importantly, CoSMoS identified an additional 18 candidate peroxisomal proteins and 300 candidate oleate-regulated genes. Preliminary localization studies suggest that these include at least 10 previously unknown peroxisomal proteins. Phenotypic studies of selected gene disruption mutants suggests that several of these new peroxisomal proteins play roles in growth on fatty acids, one is involved in peroxisome biogenesis and at least two are required for synthesis of lysine, a heretofore unrecognized role for peroxisomes. These results expand our understanding of peroxisome content and function, demonstrate the utility of CoSMoS for context-sensitive motif scanning, and point to the benefits of improved in silico genome analysis.

Gene Expression Regulation, Fungal↗

ACS4, a primary indoleacetic acid-responsive gene encoding 1-aminocyclopropane-1-carboxylate synthase in Arabidopsis thaliana. Structural characterization, expression in Escherichia coli, and expression characteristics in response to auxin [corrected].

1-Aminocyclopropane-1-carboxylic acid (ACC) synthase is the key regulatory enzyme in the biosynthetic pathway of the plant hormone ethylene. The enzyme is encoded by a divergent multigene family in Arabidopsis thaliana, comprising at least five genes, ACS1-5 (Liang, X., Abel, S., Keller, J.A., Shen,N. N.F., and Theologis, A. (1992) Poc. Natl. Acad. Sci. U.S.A. 89, 11046-11050). In etiolated seedlings, ACS4 is specifically induced by indoleacetic acid (IAA). The response to IAA is rapid (within 25 min) and insensitive to protein synthesis inhibition, suggesting that the ACS4 gene expression is a primary response to IAA. The ACS4 mRNA accumulation displays a biphasic dose-response curve which is optimal at 10 microM of IAA. However, IAA concentrations as low as 100 microM are sufficient to enhance the basal level of ACS4 mRNA. The expression of ACS4 is defective in the Arabidopsis auxin-resistant mutant lines axr1-12, axr2-1, and aux1-7. ACS4 mRNA levels are severely reduced in axr1-12 and axr2-1 but are only 1.5-fold lower in aux1-7. IAA inducibility is abolished in axr2-1. The ACS4 gene was isolated and structurally characterized. The promoter contains four sequence motifs reminiscent of functionally defined auxin-responsive cis-elements in the early auxin-inducible genes PS-IAA4/5 from pea and GH3 from soybean. Conceptual translation of the coding region predicts a protein with a molecular mass of 53,795 Da and a theoretical isoelectric point of 8.2. The ACS4 polypeptide contains the 11 invariant amino acid residues conserved between aminotransferases and ACC synthases from various plant species. An ACS4 cDNA was generated by reverse transcriptase-polymerase chain reaction, and the authenticity was confirmed by expression of ACC synthase activity in Escherichia coli.

Amino Acid Sequence↗

Promoter structure and transcriptional activation of the murine TSG-14 gene encoding a tumor necrosis factor/interleukin-1-inducible pentraxin protein.

Human TNF-stimulated gene 14 (TSG-14) encodes a secreted 42-kDa glycoprotein that shows significant homology to proteins of the pentraxin family, which includes the acute phase reactants C-reactive protein and serum amyloid P component. Levels of TSG-14 protein (also termed PTX-3) become elevated in the serum of mice and humans after injection with bacterial lipopolysaccharide, but in contrast to conventional acute phase proteins, the bulk of TSG-14 synthesis in the intact organism occurs outside the liver. In the present study we cloned and partially sequenced murine genomic TSG-14 DNA. Analysis of the coding region predicts a high degree of amino acid sequence homology between murine and human TSG-14 (88 and 75% identity in the first and second exons, respectively). The promoter of the TSG-14 gene lacks consensus sequences for either a TATA box or CCAAT box. Primer extension analysis and S1 nuclease protection assay revealed one major transcription start site, situated within a consensus sequence for an initiator element. Sequence analysis of a approximately 1.4-kilobase pair fragment of the 5'-flanking region of the TSG-14 gene revealed the presence of numerous potential enhancer binding elements, including six NF-IL6-like sites, four AP-1, one AP-2, one NF-kB, two Sp1, two interferon-gamma-activated sites (GAS), one Hox-1.3, and five binding sites for Ets family members. Transfection of BALB/c 3T3 cells with promoter DNA fragments linked to the luciferase reporter gene revealed that the 5'-flanking region of the TSG-14 gene comprises elements that can mediate a basal level of transcription and inducibility by TNF.

Amino Acid Sequence↗

AP1-mediated multidrug resistance in Saccharomyces cerevisiae requires FLR1 encoding a transporter of the major facilitator superfamily.

We have isolated a Candida albicans gene that confers resistance to the azole derivative fluconazole (FCZ) when overexpressed in Saccharomyces cerevisiae. This gene encodes a protein highly homologous to S. cerevisiae yAP-1, a bZip transcription factor known to mediate cellular resistance to toxicants such as cycloheximide (CYH), 4-nitroquinoline N-oxide (4-NQO), cadmium, and hydrogen peroxide. The gene was named CAP1, for C. albicans AP-1. Cap1 and yAP-1 are functional homologues, since CAP1 expression in a yap1 mutant strain partially restores the ability of the cells to grow on toxic concentrations of cadmium or hydrogen peroxide. We have found that the expression of YBR008c, an open reading frame identified in the yeast genome sequencing project and predicted to code for a multidrug transporter of the major facilitator superfamily, is dramatically induced in S. cerevisiae cells overexpressing CAP1. Overexpression of either CAP1 or YAP1 in a wild-type strain results in resistance to FCZ, CYH, and 4-NQO, whereas such resistance is completely abrogated (FCZ and CYH) or strongly reduced (4-NQO) in a ybr008c deletion mutant, demonstrating that YBR008c is involved in YAP1- and CAP1-mediated multidrug resistance. YBR008c has been renamed FLR1, for fluconazole resistance 1. The expression of an FLR1-lacZ reporter construct is strongly induced by the overexpression of either CAP1 or YAP1, indicating that the FLR1 gene is transcriptionally regulated by the Cap1 and yAP-1 proteins. Taken collectively, our results demonstrate that FLR1 represents a new YAP1-controlled multidrug resistance molecular determinant in S. cerevisiae. A similar detoxification pathway is also likely to operate in C. albicans.

Amino Acid Sequence↗

A 29-kilodalton Golgi soluble N-ethylmaleimide-sensitive factor attachment protein receptor (Vti1-rp2) implicated in protein trafficking in the secretory pathway.

Expressed sequence tags coding for a potential SNARE (soluble N-ethylmaleimide-sensitive factor attachment protein receptor) were revealed during data base searches. The deduced amino acid sequence of the complete coding region predicts a 217-residue protein with a COOH-terminal hydrophobic membrane anchor. Affinity-purified antibodies raised against the cytoplasmic region of this protein specifically detect a 29-kilodalton integral membrane protein enriched in the Golgi membrane. Indirect immunofluorescence microscopy reveals that this protein is mainly associated with the Golgi apparatus. When detergent extracts of the Golgi membrane are incubated with immobilized glutathione S-transferase alpha soluble N-ethylmaleimide-sensitive factor attachment protein (GST-alpha-SNAP), this protein was specifically retained. This protein has been independently identified and termed Vti1-rp2, and it is homologous to Vti1p, a yeast Golgi SNARE. We further show that Vti1-rp2 can be qualitatively coimmunoprecipitated with Golgi syntaxin 5 and syntaxin 6, suggesting that Vti1-rp2 exists in at least two distinct Golgi SNARE complexes. In cells microinjected with antibodies against Vti1-rp2, transport of the envelope protein (G-protein) of vesicular stomatitis virus from the endoplasmic reticulum to the plasma membrane was specifically arrested at the Golgi apparatus, providing further evidence for functional importance of Vti1-rp2 in protein trafficking in the secretory pathway.

Amino Acid Sequence↗

A family of yeast proteins mediating bidirectional vacuolar amino acid transport.

Seven genes in Saccharomyces cerevisiae are predicted to code for membrane-spanning proteins (designated AVT1-7) that are related to the neuronal gamma-aminobutyric acid-glycine vesicular transporters. We have now demonstrated that four of these proteins mediate amino acid transport in vacuoles. One protein, AVT1, is required for the vacuolar uptake of large neutral amino acids including tyrosine, glutamine, asparagine, isoleucine, and leucine. Three proteins, AVT3, AVT4, and AVT6, are involved in amino acid efflux from the vacuole and, as such, are the first to be shown directly to transport compounds from the lumen of an acidic intracellular organelle. This function is consistent with the role of the vacuole in protein degradation, whereby accumulated amino acids are exported to the cytosol. Protein AVT6 is responsible for the efflux of aspartate and glutamate, an activity that would account for their exclusion from vacuoles in vivo. Transport by AVT1 and AVT6 requires ATP for function and is abolished in the presence of nigericin, indicating that the same pH gradient can drive amino acid transport in opposing directions. Efflux of tyrosine and other large neutral amino acids by the two closely related proteins, AVT3 and AVT4, is similar in terms of substrate specificity to transport system h described in mammalian lysosomes and melanosomes. These findings suggest that yeast AVT transporter function has been conserved to control amino acid flux in vacuolar-like organelles.

Amino Acid Sequence↗

Molecular cloning and characterization of STAMP1, a highly prostate-specific six transmembrane protein that is overexpressed in prostate cancer.

We have identified a novel gene, six transmembrane protein of prostate 1 (STAMP1), which is largely specific to prostate for expression and is predicted to code for a 490-amino acid six transmembrane protein. Using a form of STAMP1 labeled with green fluorescent protein in quantitative time-lapse and immunofluorescence confocal microscopy, we show that STAMP1 is localized to the Golgi complex, predominantly to the trans-Golgi network, and to the plasma membrane. STAMP1 also localizes to vesicular tubular structures in the cytosol and colocalizes with the early endosome antigen 1 (EEA1), suggesting that it may be involved in the secretory/endocytic pathways. STAMP1 is highly expressed in the androgen-sensitive, androgen receptor-positive prostate cancer cell line LNCaP, but not in androgen receptor-negative prostate cancer cell lines PC-3 and DU145. Furthermore, STAMP1 expression is significantly lower in the androgen-dependent human prostate xenograft CWR22 compared with the relapsed derivative CWR22R, suggesting that its expression may be deregulated during prostate cancer progression. Consistent with this notion, in situ analysis of human prostate cancer specimens indicated that STAMP1 is expressed exclusively in the epithelial cells of the prostate and its expression is significantly increased in prostate tumors compared with normal glands. Taken together, these data suggest that STAMP1 may have an important role in the normal prostate cell as well as in prostate cancer progression.

Amino Acid Sequence↗

Vibrational local modes in DNA polymer.

Where the translational symmetry of a long polymer chain is interrupted, characteristic vibrations of the molecule are possible in which only those atoms at or relatively near the defect site partake of the motion. This contrasts with the more common vibrational states in which the motion propagates along the chain as a sound wave. Examples of readily producible local defects include broken bonds, missing atoms or groups, and extra links as are found, e.g., in thymine dimers. For each different defect, the spectrum of local mode frequencies is characteristic of its structure. Hence the local modes give direct information about the nature of the defect and can serve as a diagnostic signature of the polymer chain lesion. We have developed and are using algorithms and fortran code to predict the existence and nature of local modes based on their atomic structures. We have studied examples of different defects and found their eigen-frequencies and eigenvectors. For the simplest case of a broken hydrogen bond in a single A-T unit in a long homopolymer dA.dT chain, we display stereo views of the vibrating unit side-by-side with the undisturbed molecule for the three local modes occurring below 300 cm-1 in frequency.

Algorithms↗

Transport, fate and speciation of heavy metals (Pb, Zn, Cu, Cd) in mine drainage: geochemical modeling and anodic stripping voltammetric analysis.

The maximum concentrations (ppb) of heavy metals in the mine drainage (pH: down to 3.3) of Chonam-ri creek in the abandoned Kwangyang gold-silver mine, South Korea, are 22600 Zn, 2810 Cu, 182 Cd, and 109 Pb. A small, limestone-infused retention pond, about 440 meters downstream from the waste dump, plays an important role in the removal of heavy metals: the factors of reduction for Zn, Cu, Cd, and Pb are 12, 24, 14, and 14, respectively. This is due to the pH increase (up to >5.4) accompanying adsorption onto and/or coprecipitation with Fe- and Al-hydroxides (goethite and gibbsite). From the waste dump to the pond, heavy metal concentrations also progressively decrease due to pH increase. Geochemical modeling (using the computer code WATEQ4F) predicts that free aqueous metal ions are dominant (mostly >70% for Cu and Zn, and >60% for Pb and Cd) in samples collected upstream from the pond, whereas complexing with sulfate, carbonate and hydroxyl ions becomes important in the samples collected downstream. The comparison between the concentrations of electrochemically labile species (determined by Anodic Stripping Voltammetry) and the result of computer modeling shows that Cd and Zn are present predominantly as labile inorganic species throughout the whole range of the creek. However, Cu and Pb in the samples collected downstream from the pond largely form electrochemically inert species (possibly, metal-organic complexes). The above results indicate that the retention pond is effective in reducing the toxicity of heavy metals, especially Cu and Pb.

Electrochemistry↗

A temperature-sensitive mutation of the Schizosaccharomyces pombe gene nuc2+ that encodes a nuclear scaffold-like protein blocks spindle elongation in mitotic anaphase.

A temperature-sensitive mutant nuc2-663 of the fission yeast Schizosaccharomyces pombe specifically blocks mitotic spindle elongation at restrictive temperature so that nuclei in arrested cells contain a short uniform spindle (approximately 3-micron long), which runs through a metaphase plate-like structure consisting of three condensed chromosomes. In the wild-type or in the mutant cells at permissive temperature, the spindle is fully extended approximately 15-micron long in anaphase. The nuc2' gene was cloned in a 2.4-kb genomic DNA fragment by transformation, and its complete nucleotide sequence was determined. Its coding region predicts a 665-residues internally repeating protein (76.250 mol wt). By immunoblots using anti-sera raised against lacZ-nuc2+ fused proteins, a polypeptide (designated p67; 67,000 mol wt) encoded by nuc2+ is detected in the wild-type S. pombe extracts; the amount of p67 is greatly increased when multi-copy or high-expression plasmids carrying the nuc2+ gene are introduced into the S. pombe cells. Cellular fractionation and Percoll gradient centrifugation combined with immunoblotting show that p67 cofractionates with nuclei and is enriched in resistant structure that is insoluble in 2 M NaCl, 25 mM lithium 3,5'-diiodosalicylate, and 1% Triton but is soluble in 8 M urea. In nuc2 mutant cells, however, soluble p76, perhaps an unprocessed precursor, accumulates in addition to insoluble p67. The role of nuc2+ gene may be to interconnect nuclear and cytoskeletal functions in chromosome separation.

Amino Acid Sequence↗

A retinoic acid responsive gene MK found in the teratocarcinoma system is expressed in spatially and temporally controlled manner during mouse embryogenesis.

A newly identified gene MK is transiently expressed in early stages of retinoic acid-induced differentiation of embryonal carcinoma cells (Kadomatsu, K., M. Tomomura, and T. Muramatsu, 1988. Biochem. Biophys. Res. Commun. 151:1312-1318). MK gene has been predicted to code a polypeptide that is rich in basic amino acids and cysteine and is not related to any other peptides so far reported. In the present study, we investigated MK expression during mouse embryogenesis by in situ hybridization. The MK transcript was detected all over the embryo proper of the 7-d embryo, while it was not detectable in the 5-d embryo. The ubiquitous expression continued in the 9-d embryo proper. On the 11th-13th d of gestation, the sites where MK gene was intensely expressed became progressively restricted; these sites were the brain ectoderm around the lens and brain ventricles, the anterior lobe of the pituitary gland, the upper and lower jaw, the caudal sclerotomic half of vertebral column, the limbs, the stomach, and the epithelial tissues of the lung, the pancreas, the small intestine, and the metanephros. These areas include the region where secondary embryonic induction is prominent. In the 15-d embryo, only the kidney expressed MK significantly. These data suggest that MK gene plays a fundamental role in the differentiation of a wide variety of cells; MK gene may also play some specific roles in generation of epithelial tissues, and remodeling of mesoderm.

Animals↗

Isolation, characterization, and expression of cDNAs encoding murine alpha-mannosidase II, a Golgi enzyme that controls conversion of high mannose to complex N-glycans.

Golgi alpha-mannosidase II (GlcNAc transferase I-dependent alpha 1,3[alpha 1,6] mannosidase, EC 3.2.1.114) catalyzes the final hydrolytic step in the N-glycan maturation pathway acting as the committed step in the conversion of high mannose to complex type structures. We have isolated overlapping clones from a murine cDNA library encoding the full length alpha-mannosidase II open reading frame and most of the 5' and 3' untranslated region. The coding sequence predicts a type II transmembrane protein with a short cytoplasmic tail (five amino acids), a single transmembrane domain (21 amino acids), and a large COOH-terminal catalytic domain (1,124 amino acids). This domain organization which is shared with the Golgi glycosyl-transferases suggests that the common structural motifs may have a functional role in Golgi enzyme function or localization. Three sets of polyadenylated clones were isolated extending 3' beyond the open reading frame by as much as 2,543 bp. Northern blots suggest that these polyadenylated clones totaling 6.1 kb in length correspond to minor message species smaller than the full length message. The largest and predominant message on Northern blots (7.5 kb) presumably extends another approximately 1.4-kb downstream beyond the longest of the isolated clones. Transient expression of the alpha-mannosidase II cDNA in COS cells resulted in 8-12-fold overexpression of enzyme activity, and the appearance of cross-reactive material in a perinuclear membrane array consistent with a Golgi localization. A region within the catalytic domain of the alpha-mannosidase II open reading frame bears a strong similarity to a corresponding sequence in the rat liver endoplasmic reticulum alpha-mannosidase and the vacuolar alpha-mannosidase of Saccharomyces cerevisiae. Partial human alpha-mannosidase II cDNA clones were also isolated and the gene was localized to human chromosome 5.

Amino Acid Sequence↗

Distinctly different gene structure of KLK4/KLK-L1/prostase/ARM1 compared with other members of the kallikrein family: intracellular localization, alternative cDNA forms, and Regulation by multiple hormones.

The tissue kallikreins (KLKs) form a family of serine proteases that are involved in processing of polypeptide precursors and have important roles in a variety of physiologic and pathological processes. Common features of all tissue kallikrein genes identified to date in various species include a similar genomic organization of five exons, a conserved triad of amino acids for serine protease catalytic activity, and a signal peptide sequence encoded in the first exon. Here, we show that KLK4/KLK-L1/prostase/ARM1 (hereafter called KLK4) is the first significantly divergent member of the kallikrein family. The exon predicted to code for a signal peptide is absent in KLK4, which is likely to affect the function of the encoded protein. Green fluorescent protein (GFP)-tagged KLK4 has a distinct perinuclear localization, suggesting that its primary function is inside the cell, in contrast to the other tissue kallikreins characterized so far that have major extracellular functions. There are at least two differentially spliced, truncated variants of KLK4 that are either exclusively or predominantly localized to the nucleus when labeled with GFP. Furthermore, KLK4 expression is regulated by multiple hormones in prostate cancer cells and is deregulated in the androgen-independent phase of prostate cancer. These findings demonstrate that KLK4 is a unique member of the kallikrein family that may have a role in the progression of prostate cancer.

Alternative Splicing↗

Isolation and nucleotide sequence analysis of a cloned cDNA encoding the beta-subunit of bovine follicle-stimulating hormone.

Two different cDNAs containing sequences coding for the beta-subunit of bovine follicle stimulating hormone (FSH-beta) have been isolated from a phage lambda gt11 bovine pituitary cDNA library. The complete nucleotide sequence of both clones was determined, and the combined sequence represents most of FSH-beta mRNA. The combined sequence contains 46 nucleotides of 5'-untranslated sequence followed by 387 nucleotides of coding sequence. The coding sequence predicts a 19-amino-acid amino-terminal precursor segment followed by the 110-amino-acid sequence of mature bovine FSH-beta. The cDNA sequence demonstrates the presence of a long 3'-untranslated region containing 1295 bases followed by a segment representing the poly(A) portion of the mRNA. Thus, the combined sequence of the cDNAs suggests a minimal size of 1.7 kb for FSH-beta mRNA. Analysis of FSH-beta sequences present in bovine pituitary mRNA demonstrated the presence of an mRNA with a size of about 2.0 kb. This apparent discrepancy is probably due to the presence of a several-hundred nucleotide tract of poly(A) at the 3' terminus of the mRNA. Comparison of the amino acid sequence predicted from the cDNA with the known amino acid sequence of the beta-subunit of FSH from several different species demonstrates that the protein has been highly conserved.

Amino Acid Sequence↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

A parallel neural network simulator on the connection machine CM-5.

We here present a parallel implementation of artificial neural networks on the connection machine CM-5 and compare it with other parallel implementations on SIMD and MIMD architectures. This parallel implementation was developed with the goal of efficiently training large neural networks with huge training pattern sets for applications in molecular biology, in particular the prediction of coding regions in DNA sequences. The implementation uses training pattern parallelism and makes use of the parallel I/O facilities of the CM-5 and its efficient reduction operations available within the control network to achieve a high scalability. The parallel simulator obtains a maximum speed of 149.25 MCUPS for training feedforward networks with backpropagation on a 512 processor CM-5 system without using the CM-5 vector facility. The implementation poses no restriction on the type of network topology and works with different batch training algorithms like BP. Quickprop and Rprop.

Algorithms↗

Interactive InterPro-based comparisons of proteins in whole genomes.

MOTIVATION: The SWISS-PROT group at the EBI has developed the Proteome Analysis Database utilizing existing resources and providing comprehensive and integrated comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archaea and eukaryotes. The Proteome Analysis Database is accompanied by a program that has been designed to carry out interactive InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Computational Biology↗