Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Isolation, characterization, and expression of cDNAs encoding murine alpha-mannosidase II, a Golgi enzyme that controls conversion of high mannose to complex N-glycans.

Golgi alpha-mannosidase II (GlcNAc transferase I-dependent alpha 1,3[alpha 1,6] mannosidase, EC 3.2.1.114) catalyzes the final hydrolytic step in the N-glycan maturation pathway acting as the committed step in the conversion of high mannose to complex type structures. We have isolated overlapping clones from a murine cDNA library encoding the full length alpha-mannosidase II open reading frame and most of the 5' and 3' untranslated region. The coding sequence predicts a type II transmembrane protein with a short cytoplasmic tail (five amino acids), a single transmembrane domain (21 amino acids), and a large COOH-terminal catalytic domain (1,124 amino acids). This domain organization which is shared with the Golgi glycosyl-transferases suggests that the common structural motifs may have a functional role in Golgi enzyme function or localization. Three sets of polyadenylated clones were isolated extending 3' beyond the open reading frame by as much as 2,543 bp. Northern blots suggest that these polyadenylated clones totaling 6.1 kb in length correspond to minor message species smaller than the full length message. The largest and predominant message on Northern blots (7.5 kb) presumably extends another approximately 1.4-kb downstream beyond the longest of the isolated clones. Transient expression of the alpha-mannosidase II cDNA in COS cells resulted in 8-12-fold overexpression of enzyme activity, and the appearance of cross-reactive material in a perinuclear membrane array consistent with a Golgi localization. A region within the catalytic domain of the alpha-mannosidase II open reading frame bears a strong similarity to a corresponding sequence in the rat liver endoplasmic reticulum alpha-mannosidase and the vacuolar alpha-mannosidase of Saccharomyces cerevisiae. Partial human alpha-mannosidase II cDNA clones were also isolated and the gene was localized to human chromosome 5.

Amino Acid Sequence↗

Distinctly different gene structure of KLK4/KLK-L1/prostase/ARM1 compared with other members of the kallikrein family: intracellular localization, alternative cDNA forms, and Regulation by multiple hormones.

The tissue kallikreins (KLKs) form a family of serine proteases that are involved in processing of polypeptide precursors and have important roles in a variety of physiologic and pathological processes. Common features of all tissue kallikrein genes identified to date in various species include a similar genomic organization of five exons, a conserved triad of amino acids for serine protease catalytic activity, and a signal peptide sequence encoded in the first exon. Here, we show that KLK4/KLK-L1/prostase/ARM1 (hereafter called KLK4) is the first significantly divergent member of the kallikrein family. The exon predicted to code for a signal peptide is absent in KLK4, which is likely to affect the function of the encoded protein. Green fluorescent protein (GFP)-tagged KLK4 has a distinct perinuclear localization, suggesting that its primary function is inside the cell, in contrast to the other tissue kallikreins characterized so far that have major extracellular functions. There are at least two differentially spliced, truncated variants of KLK4 that are either exclusively or predominantly localized to the nucleus when labeled with GFP. Furthermore, KLK4 expression is regulated by multiple hormones in prostate cancer cells and is deregulated in the androgen-independent phase of prostate cancer. These findings demonstrate that KLK4 is a unique member of the kallikrein family that may have a role in the progression of prostate cancer.

Alternative Splicing↗

Isolation and nucleotide sequence analysis of a cloned cDNA encoding the beta-subunit of bovine follicle-stimulating hormone.

Two different cDNAs containing sequences coding for the beta-subunit of bovine follicle stimulating hormone (FSH-beta) have been isolated from a phage lambda gt11 bovine pituitary cDNA library. The complete nucleotide sequence of both clones was determined, and the combined sequence represents most of FSH-beta mRNA. The combined sequence contains 46 nucleotides of 5'-untranslated sequence followed by 387 nucleotides of coding sequence. The coding sequence predicts a 19-amino-acid amino-terminal precursor segment followed by the 110-amino-acid sequence of mature bovine FSH-beta. The cDNA sequence demonstrates the presence of a long 3'-untranslated region containing 1295 bases followed by a segment representing the poly(A) portion of the mRNA. Thus, the combined sequence of the cDNAs suggests a minimal size of 1.7 kb for FSH-beta mRNA. Analysis of FSH-beta sequences present in bovine pituitary mRNA demonstrated the presence of an mRNA with a size of about 2.0 kb. This apparent discrepancy is probably due to the presence of a several-hundred nucleotide tract of poly(A) at the 3' terminus of the mRNA. Comparison of the amino acid sequence predicted from the cDNA with the known amino acid sequence of the beta-subunit of FSH from several different species demonstrates that the protein has been highly conserved.

Amino Acid Sequence↗

Analysis of particle deposition in the turbinate and olfactory regions using a human nasal computational fluid dynamics model.

The human nasal passages effectively filter particles from inhaled air. This prevents harmful pollutants from reaching susceptible pulmonary airways, but may leave the nasal mucosa vulnerable to potentially injurious effects from inhaled toxicants. This filtering property may also be strategically used for aerosolized nasal drug delivery. The nasal route has recently been considered as a means of delivering systemically acting drugs due to the large absorptive surface area available in close proximity to the nostrils. In this study, a computational fluid dynamics (CFD) model of nasal airflow was used with a particle transport and deposition code to predict localized deposition of inhaled particles in human nasal passages. The model geometry was formed from MRI scan tracings of the nasal passages of a healthy adult male. Spherical particles ranging in size from 5 to 50 microm were released from the nostrils. Particle trajectories and deposition sites were calculated in the presence of steady-state inspiratory airflow at volumetric flow rates of 7.5, 15, and 30 L/min. The nasal valve, turbinates, and olfactory region were defined in the CFD model so that particles depositing in these regions could be identified and correlated with their release positions on the nostril surfaces. When plotted against impaction parameter, deposition efficiencies in these regions exhibited maximum values of 53%, 20%, and 3%, respectively. Analysis of preferential deposition patterns and nostril release positions under natural breathing scenarios can be used to determine optimal particle size and flow rate combinations to selectively target drug particles to specific regions of the nose.

Administration, Inhalation↗

Comparative genomics tools applied to bioterrorism defence.

Rapid advances in the genomic sequencing of bacteria and viruses over the past few years have made it possible to consider sequencing the genomes of all pathogens that affect humans and the crops and livestock upon which our lives depend. Recent events make it imperative that full genome sequencing be accomplished as soon as possible for pathogens that could be used as weapons of mass destruction or disruption. This sequence information must be exploited to provide rapid and accurate diagnostics to identify pathogens and distinguish them from harmless near-neighbours and hoaxes. The Chem-Bio Non-Proliferation (CBNP) programme of the US Department of Energy (DOE) began a large-scale effort of pathogen detection in early 2000 when it was announced that the DOE would be providing bio-security at the 2002 Winter Olympic Games in Salt Lake City, Utah. Our team at the Lawrence Livermore National Lab (LLNL) was given the task of developing reliable and validated assays for a number of the most likely bioterrorist agents. The short timeline led us to devise a novel system that utilised whole-genome comparison methods to rapidly focus on parts of the pathogen genomes that had a high probability of being unique. Assays developed with this approach have been validated by the Centers for Disease Control (CDC). They were used at the 2002 Winter Olympics, have entered the public health system, and have been in continual use for non-publicised aspects of homeland defence since autumn 2001. Assays have been developed for all major threat list agents for which adequate genomic sequence is available, as well as for other pathogens requested by various government agencies. Collaborations with comparative genomics algorithm developers have enabled our LLNL team to make major advances in pathogen detection, since many of the existing tools simply did not scale well enough to be of practical use for this application. It is hoped that a discussion of a real-life practical application of comparative genomics algorithms may help spur algorithm developers to tackle some of the many remaining problems that need to be addressed. Solutions to these problems will advance a wide range of biological disciplines, only one of which is pathogen detection. For example, exploration in evolution and phylogenetics, annotating gene coding regions, predicting and understanding gene function and regulation, and untangling gene networks all rely on tools for aligning multiple sequences, detecting gene rearrangements and duplications, and visualising genomic data. Two key problems currently needing improved solutions are: (1) aligning incomplete, fragmentary sequence (eg draft genome contigs or arbitrary genome regions) with both complete genomes and other fragmentary sequences; and (2) ordering, aligning and visualising non-colinear gene rearrangements and inversions in addition to the colinear alignments handled by current tools.

Amino Acid Sequence↗

A parallel neural network simulator on the connection machine CM-5.

We here present a parallel implementation of artificial neural networks on the connection machine CM-5 and compare it with other parallel implementations on SIMD and MIMD architectures. This parallel implementation was developed with the goal of efficiently training large neural networks with huge training pattern sets for applications in molecular biology, in particular the prediction of coding regions in DNA sequences. The implementation uses training pattern parallelism and makes use of the parallel I/O facilities of the CM-5 and its efficient reduction operations available within the control network to achieve a high scalability. The parallel simulator obtains a maximum speed of 149.25 MCUPS for training feedforward networks with backpropagation on a 512 processor CM-5 system without using the CM-5 vector facility. The implementation poses no restriction on the type of network topology and works with different batch training algorithms like BP. Quickprop and Rprop.

Algorithms↗

Interactive InterPro-based comparisons of proteins in whole genomes.

MOTIVATION: The SWISS-PROT group at the EBI has developed the Proteome Analysis Database utilizing existing resources and providing comprehensive and integrated comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archaea and eukaryotes. The Proteome Analysis Database is accompanied by a program that has been designed to carry out interactive InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Computational Biology↗

Identification and expression analysis of Drosophila melanogaster genes encoding beta-hexosaminidases of the sperm plasma membrane.

Sperm surface beta-N-acetylhexosaminidases are among the molecules mediating early gamete interactions in invertebrates and vertebrates, including man. The plasma membrane of Drosophila spermatozoa contains two beta-N-acetylhexosaminidases, DmHEXA and DmHEXB, which are required for egg fertilization. Here, we demonstrate that three putative Drosophila melanogaster genes predicted to code for beta-N-acetylhexosaminidases, Hexo1, Hexo2, and fdl, are all expressed in the male germ line. fdl codes for a homolog of the alpha-subunit of the mammalian lysosomal beta-N-acetylhexosaminidase Hex A. Hexo1 and Hexo2 encode two homologs of the beta-subunit of all known beta-N-acetylhexosaminidases, which we have named beta(1) and beta(2), respectively. Immunoblot analysis of sperm proteins indicated that the gene products associate in different heterodimeric combinations forming DmHEXA, with an alphabeta(2) structure, and DmHEXB, with a beta(1)beta(2) structure. Immunofluorescence demonstrated that all the gene products localized to the sperm plasma membrane. Although none of the genes was testis-specific, fdl was highly and preferentially expressed in the testis, whereas Hexo1 and Hexo2 showed broader tissue expression. Enzyme assays carried out on testis and on a variety of somatic tissues corroborated the results of gene expression analysis. These findings for the first time show the in vivo expression in insects of genes encoding beta-N-acetylhexosaminidases, the only molecules so far identified as involved in sperm/egg recognition in this class, whereas in mammals, the organisms where these enzymes have been best studied, only two types of polypeptide chains forming dimeric functional beta-N-acetylhexosaminidases are present in Drosophila three different gene products are available that might generate numerous dimeric isoforms.

Amino Acid Sequence↗

Aprataxin, a novel protein that protects against genotoxic stress.

Ataxia-oculomotor apraxia (AOA1) is a neurological disorder with symptoms that overlap those of ataxia-telangiectasia, a syndrome characterized by abnormal responses to double-strand DNA breaks and genome instability. The gene mutated in AOA1, APTX, is predicted to code for a protein called aprataxin that contains domains of homology with proteins involved in DNA damage signalling and repair. We demonstrate that aprataxin is a nuclear protein, present in both the nucleoplasm and the nucleolus. Mutations in the APTX gene destabilize the aprataxin protein, and fusion constructs of enhanced green fluorescent protein and aprataxin, representing deletions of putative functional domains, generate highly unstable products. Cells from AOA1 patients are characterized by enhanced sensitivity to agents that cause single-strand breaks in DNA but there is no evidence for a gross defect in single-strand break repair. Sensitivity to hydrogen peroxide and the resulting genome instability are corrected by transfection with full-length aprataxin cDNA. We also demonstrate that aprataxin interacts with the repair proteins XRCC1, PARP-1 and p53 and that it co-localizes with XRCC1 along charged particle tracks on chromatin. These results demonstrate that aprataxin influences the cellular response to genotoxic stress very likely by its capacity to interact with a number of proteins involved in DNA repair.

Apraxias↗

A new rule for analyzing homologous coding sequences in DNA.

A new rule is proposed for detecting homology between DNA sequences coding for proteins. Simple coding considerations predict that if two DNA sequences are homologous because of a common ancestry, they should share sequence similarities primarily in the same translation phase, with their codons aligned. Similarities which are in phase are shown to be more frequent between related sequences than between unrelated or random ones. But similar segments which are out of translation phase are no more frequent between related than unrelated sequences. Duplication and concatenation of genetic elements, followed by gene duplication and random mutations would lead to the patterns observed. Similarities are examined between various immunologically important sequences, including immunoglobulins, MHC products, and the T lymphocyte antigen Thy 1.

Antigens, Surface↗

Complete nucleotide sequence of tobacco streak virus RNA 3.

Double-stranded cDNA of in vitro polyadenylated tobacco streak virus (TSV) RNA 3 has been cloned and sequenced. The complete primary structure of 2,205 nucleotides reveals two open reading frames flanked by a leader sequence of 210 bases, an intercistronic region of 123 nucleotides and a 3'-extracistronic sequence of 288 nucleotides. The 5'-terminal open reading frame codes for a Mr 31,742 protein, which probably corresponds to the only in vitro translation product of TSV RNA 3. The 3'-terminal coding region predicts a Mr 26,346 protein, probably the viral coat protein, which is the translation product of the subgenomic messenger, RNA 4. Although the coat proteins of alfalfa mosaic virus (A1MV) and TSV are functionally equivalent in activating their own and each others genomes, no homology between the primary structures of those two proteins is detectable.

Amino Acid Sequence↗

Nucleotide sequence of a gene from chromosome 1D of wheat encoding a HMW-glutenin subunit.

A high molecular weight glutenin gene in hexaploid wheat has been isolated by cloning in bacteriophage lambda and characterized. The gene corresponds to polypeptide 12 encoded by chromosome 1D in the variety "Chinese Spring". The coding sequence predicted contains seven cysteine residues six of which flank a central repetitive region comprising more than 70% of the polypeptide. These findings are related to the role of high molecular weight subunits in the viscoelastic theory of gluten structure.

Amino Acid Sequence↗

DNA sequence of the herpes simplex virus type 1 gene encoding glycoprotein gH, and identification of homologues in the genomes of varicella-zoster virus and Epstein-Barr virus.

We have determined the sequence of herpes simplex virus type 1 DNA around the previously mapped location of sequences encoding an epitope of glycoprotein gH, and have deduced the structure of the gH gene and the amino acid sequence of gH. The unprocessed polypeptide is predicted to contain 838 amino acids, and to possess an N-terminal signal sequence and a C-terminal transmembrane sequence. Temperature-sensitive mutant tsQ26 maps within the predicted gH coding sequence. Homologous genes were identified in the genomes of two other herpesviruses, namely varicella-zoster virus and Epstein-Barr virus.

Amino Acid Sequence↗

Isolation and characterization of the human tyrosine aminotransferase gene.

Structure and sequence of the human gene for tyrosine aminotransferase (TAT) was determined by analysis of cDNA and genomic clones. The gene extends over 10.9 kbl and consists of 12 exons giving rise to a 2,754 nucleotide long mRNA (excluding the poly(A)tail). The human TAT gene is predicted to code for a 454 amino acid protein of molecular weight 50,399 dalton. The overall sequence identity within the coding region of the human and the previously characterized rat TAT genes is 87% at the nucleotide and 92% at the protein level. A minor human TAT mRNA results from the use of an alternative polyadenylation signal in the 3' exon which is present but not used at the corresponding position in the rat TAT gene. The non-coding region of the 3' exon contains a complete Alu element which is absent in the rat TAT gene but present in apes and old world monkeys. Two functional glucocorticoid response elements (GREs) reside 2.5 kb upstream of the rat TAT gene. The DNA sequence of the corresponding region of the human TAT gene shows the distal GRE mutated and the proximal GRE replaced by Alu elements.

Amino Acid Sequence↗

Complete genome sequence of the alkaliphilic bacterium Bacillus halodurans and genomic sequence comparison with Bacillus subtilis.

The 4 202 353 bp genome of the alkaliphilic bacterium Bacillus halodurans C-125 contains 4066 predicted protein coding sequences (CDSs), 2141 (52.7%) of which have functional assignments, 1182 (29%) of which are conserved CDSs with unknown function and 743 (18. 3%) of which have no match to any protein database. Among the total CDSs, 8.8% match sequences of proteins found only in Bacillus subtilis and 66.7% are widely conserved in comparison with the proteins of various organisms, including B.subtilis. The B. halodurans genome contains 112 transposase genes, indicating that transposases have played an important evolutionary role in horizontal gene transfer and also in internal genetic rearrangement in the genome. Strain C-125 lacks some of the necessary genes for competence, such as comS, srfA and rapC, supporting the fact that competence has not been demonstrated experimentally in C-125. There is no paralog of tupA, encoding teichuronopeptide, which contributes to alkaliphily, in the C-125 genome and an ortholog of tupA cannot be found in the B.subtilis genome. Out of 11 sigma factors which belong to the extracytoplasmic function family, 10 are unique to B. halodurans, suggesting that they may have a role in the special mechanism of adaptation to an alkaline environment.

ATP-Binding Cassette Transporters↗

Proteome Analysis Database: online application of InterPro and CluSTr for the functional classification of proteins in whole genomes.

The SWISS-PROT group at EBI has developed the Proteome Analysis Database utilising existing resources and providing comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archaea and eukaryotes (http://www.ebi.ac. uk/proteome/). The two main projects used, InterPro and CluSTr, give a new perspective on families, domains and sites and cover 31-67% (InterPro statistics) of the proteins from each of the complete genomes. CluSTr covers the three complete eukaryotic genomes and the incomplete human genome data. The Proteome Analysis Database is accompanied by a program that has been designed to carry out InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Animals↗

The Proteome Analysis database: a tool for the in silico analysis of whole proteomes.

The Proteome Analysis database (http://www.ebi.ac.uk/proteome/) has been developed by the Sequence Database Group at EBI utilizing existing resources and providing comparative analysis of the predicted protein coding sequences of the complete genomes of bacteria, archeae and eukaryotes. Three main projects are used, InterPro, CluSTr and GO Slim, to give an overview on families, domains, sites, and functions of the proteins from each of the complete genomes. Complete proteome analysis is available for a total of 89 proteome sets. A specifically designed application enables InterPro proteome comparisons for any one proteome against any other one or more of the proteomes in the database.

Animals↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗