Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Predictive coding”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Sequence and organization of the human mitochondrial genome.

The complete sequence of the 16,569-base pair human mitochondrial genome is presented. The genes for the 12S and 16S rRNAs, 22 tRNAs, cytochrome c oxidase subunits I, II and III, ATPase subunit 6, cytochrome b and eight other predicted protein coding genes have been located. The sequence shows extreme economy in that the genes have none or only a few noncoding bases between them, and in many cases the termination codons are not coded in the DNA but are created post-transcriptionally by polyadenylation of the mRNAs.

Base Sequence↗

Phase variation in Bordetella pertussis by frameshift mutation in a gene for a novel two-component system.

Bordetella pertussis, the aetiological agent of whooping cough, coordinately regulates the expression of many virulence-associated determinants, including filamentous haemagglutinin, pertussis toxin, adenylyl cyclase toxin, dermonecrotic toxin and haemolysin. The coordinate regulation is apparent in the repression of synthesis of these determinants in response to environmental stimuli; a phenomenon known as antigenic or phenotypic modulation. B. pertussis also varies between metastable genetic states, or phases. There is a virulent phase in which virulence-associated determinants are synthesized, and an avirulent phase in which they are not. Previous studies have shown that a genetic locus, vir, is required for expression from many virulence-associated loci, and that replacing the cloned vir locus in trans can restore the virulent phase phenotype to spontaneously occurring avirulent phase strains. Here, we show that phase variation in one series of strains is due to a frameshift mutation within an open reading frame that is predicted to code for a Vir protein product. The deduced protein sequence is similar to both components of the 'two-component' regulatory system which control gene expression in response to environmental stimuli in a range of bacterial species.

Base Sequence↗

Identification of the spinocerebellar ataxia type 2 gene using a direct identification of repeat expansion and cloning technique, DIRECT.

Spinocerebellar ataxia type 2 (SCA2) is an autosomal dominant, neurodegenerative disorder that affects the cerebellum and other areas of the central nervous system. We have devised a novel strategy, the direct identification of repeat expansion and cloning technique (DIRECT), which allows selective detection of expanded CAG repeats and cloning of the genes involved. By applying DIRECT, we identified an expanded CAG repeat of the gene for SCA2. CAG repeats of normal alleles range in size from 15 to 24 repeat units, while those of SCA2 chromosomes are expanded to 35 to 59 repeat units. The SCA2 cDNA is predicted to code for 1,313 amino acids-with the CAG repeats coding for a polyglutamine tract. DIRECT is a robust strategy for identification of pathologically expanded trinucleotide repeats and will dramatically accelerate the search for causative genes of neuropsychiatric diseases caused by trinucleotide repeat expansions.

Amino Acid Sequence↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

Functional characterization and chromosomal localization of a cloned taurine transporter from human placenta.

A cDNA clone highly related to the rat brain taurine transporter has been isolated from a human placental cDNA library. Transfection of this cDNA into HeLa cells results in a marked elevation of taurine transport activity. The activity of the cDNA-induced transporter is dependent on the presence of Na+ as well as Cl-. The Na+/Cl-/taurine stoichiometry for the cloned transporter is 2:1:1. The transporter is specific for taurine and other beta-amino acids, including beta-alanine, and exhibits high affinity for taurine (Michaelis-Menten constant approximately 6 microM). The clone consists of a coding region 1863 bp long (including the termination codon), flanked by a 376 bp-long 5' non-coding region and a 625 bp-long 3' non-coding region. The nucleotide sequence of the coding region predicts a 620-amino acid protein with a calculated M(r) of 69,853. Northern-blot analysis of poly(A)+ RNA from several human tissues indicates a complex expression pattern differing across tissues. The principal transcript, 6.9 kb in size, is expressed abundantly in placenta and skeletal muscle, at intermediate levels in heart, brain, lung, kidney and pancreas and at low levels in liver. Cultured human cell lines derived from placenta (JAR and BeWo), intestine (HT-29), cervix (HeLa) and retinal pigment epithelium (HRPE), which are known to possess Na(+)- and Cl(-)-coupled taurine transport activity, also contain the 6.9 kb transcript. Somatic cell hybrid and in situ hybridization studies indicate that the cloned taurine transporter is localized to human chromosome 3 p24-->p26.

Amino Acid Sequence↗

Cloning and expression of a novel juvenile hormone-metabolizing epoxide hydrolase during larval-pupal metamorphosis of the cabbage looper, Trichoplusia ni.

A full-length cDNA encoding for a microsomal juvenile hormone (JH)-metabolizing epoxide hydrolase (TmEH-1) was isolated from a cDNA library constructed from fat body of last stadium (wandering) cabbage loopers, Trichoplusia ni, at the exact developmental time of maximum JH epoxide hydrolase activity. TmEH-1 was 1887 base pairs in length with a 1389 base pair open reading frame encoding 463 amino acids. Amino acid sequence analysis showed that TmEH-1 was most similar to and contained the exact catalytic triad (Asp-226, Glu-403 and His-430) found in microsomal epoxide hydrolases. TmEH-1-specific message was present along with JH III epoxide hydrolase activity in fat body in feeding (days 1 and 2) and wandering (day 3) larvae with the peak in message level preceding the peak in JH epoxide hydrolase activity by 1 day. When TmEH-1 was expressed in baculovirus-infected Spodoptera frugiperda cells, a 46,000 molecular weight protein appeared on SDS-PAGE which corresponded to the predicted size coded by the TmEH-1 message and which was positively correlated with increases in JH III epoxide hydrolase activity above that of wild-type controls. In subcellular distribution studies, 58% of the juvenile hormone III epoxide hydrolase activity was in the insoluble fractions. Baculovirus expressed TmEH-1 demonstrated a higher specific activity for JH III as compared to the general EH substrates, cis- and trans-stilbene oxide. Southern blot analyses suggested that multiple epoxide hydrolase genes are present in T. ni.

Amino Acid Sequence↗

The Brucella abortus Lon functions as a generalized stress response protease and is required for wild-type virulence in BALB/c mice.

The gene encoding a Lon protease homologue has been cloned from Brucella abortus. The putative Brucella abortus Lon shares > 60% amino acid identity with its Escherichia coli counterpart and the recombinant form of this protein restores the capacity of an Escherichia coli lon mutant to resist killing by ultraviolet irradiation and regulate the expression of a cpsB:lacZ fusion to wild-type levels. A sigma32 type promoter was identified upstream of the predicted lon coding region and Northern analysis revealed that transcription of the native Brucella abortus lon increases in response to heat shock and other environmental stresses. ATP-dependent proteolytic activity was also demonstrated for purified recombinant Lon. To evaluate the capacity of the Brucella abortus Lon homologue to function as a stress response protease, the majority of the lon coding region was removed from virulent strain Brucella abortus 2308 via allelic exchange. In contrast to the parent strain, the Brucella abortus lon mutant, designated GR106, was impaired in its capacity to form isolated colonies on solid medium at 41 degrees C and displayed an increased sensitivity to killing by puromycin and H2O2. GR106 also displayed reduced survival in cultured murine macrophages and significant attenuation in BALB/c mice at 1 week post infection compared with the virulent parental strain. Beginning at 2 weeks and continuing for 6 weeks post infection, however, GR106 and 2308 displayed equivalent spleen and liver colonization levels in mice. These findings suggest that the Brucella abortus Lon homologue functions as a stress response protease that is required for wild-type virulence during the initial stages of infection in the mouse model, but is not essential for the establishment and maintenance of chronic infection in this host.

ATP-Dependent Proteases↗

Identification of two Ikaros-like transcription factors in lamprey.

The jawless Agnatha (lampreys and hagfishes) represent the phylogenetically oldest order of vertebrates that are believed to lack the adaptive immune system of jawed vertebrates. In order to search for molecular markers specific for cellular components of the adaptive immune system in lampreys, we used the polymerase chain reaction (PCR) to identify genes for transcription factors of the Ikaros family in genomic DNA and cDNA libraries from two species of lampreys, Petromyzon marinus and Lampetra fluviatilis. The mammalian Ikaros-like family of transcription factors consists of five members, Ikaros, Helios, Aiolos, Eos and Pegasus, of which the first three appear to be essential for lymphocyte development. Two different Ikaros-like genes, named IKLF1 and IKLF2, were identified in lamprey. They both have the conserved exon-intron structure of seven exons and show alternative splicing like their counterparts in jawed vertebrates. The genes code for predicted proteins of 589 and 513 amino acid residues, respectively. The proteins contain six highly conserved zinc finger motifs that are 83-91% identical to the mammalian members of the Ikaros-like family. The remaining parts of the sequences are, however, mostly unalignable. Phylogenetic analysis based on the alignable segments of the sequences does not identify the orthologous gene in jawed vertebrates but rather shows equidistance of the lamprey Ikaros-like factors to each other and to Ikaros, Helios, Aiolos and Eos. Expression studies by reverse transcription (RT)-PCR and in situ hybridization (ISH), however, provide evidence for moderate expression in presumed lymphoid tissues like the gut epithelium and for high levels of expression in the gonads, especially in the ovary.

Amino Acid Sequence↗

A promoter trap for Chlamydomonas reinhardtii: development of a gene cloning method using 5' RACE-based probes.

A promoterless radial spoke protein RSP3 gene has been used to identify promoter regions in the genome of Chlamydomonas reinhardtii. The acceptor strain pf-14 arg7 was transformed with a linearized vector containing the ARG 7.8 gene as a selection marker and a promoterless RSP3 gene. The frequency at which the motility was restored in transformants varied from 2-3%. Several of these were motile only in ammonium-free medium, indicating that the procedure could be used to select inducible promoters. Transformation of nitrogen-starved cells produced about twice as many transformants which were only motile in ammonium-free medium. Since one of the tagging vectors contained an RSP3 gene with a hybridization flag in its 3' untranslated region, it was possible to estimate the size of the new RSP3 transcripts in transformants. The results suggested that in most cases a hybrid RNA was generated consisting of the tagged gene transcript and reporter gene RNA. By 5' RACE, these parts of the new transcripts were amplified and it was shown that the generated DNA fragments could be used to clone a tagged gene. One such example, gene 2BC9, is predicted to code for a mitochondrial matrix protein. The tagging procedure will be optimized for cloning genes induced by nitrogen starvation, the cue for gametogenesis.

Amino Acid Sequence↗

Aspartyl proteases in Caenorhabditis elegans. Isolation, identification and characterization by a combined use of affinity chromatography, two-dimensional gel electrophoresis, microsequencing and databank analysis.

Crude homogenates of the nematode Caenorhabditis elegans exhibit maximal proteolytic activity under acidic pH conditions. About 90% of this activity is inhibited by the oligopeptide pepstatin, which specifically inhibits the activity of aspartyl proteases such as pepsin, cathepsins D and E or renin. We have purified enzymes responsible for this proteolytic activity by a single-step affinity chromatography on pepstatin-agarose. Analysis of the purified fraction by 1D SDS gel electrophoresis revealed six bands ranging from 35 to 52 kDa. After electrotransfer to poly(vinylidene difluoride) membranes, all bands were successfully subjected to N-terminal microsequencing. On 2D gels, the purified protein bands split into 19 spots which, after renewed microsequencing, were identified as isoelectric variants of the six proteins already described. The N-termini obtained for these proteins could be correlated to genomic DNA sequences determined in the course of the C. elegans genome sequencing project. All these sequences were predicted to code for expressed proteins as collected in the WORMPEP database. Five of the six coding sequences identified in this study were found to contain the typical active-site consensus sequence of aspartyl proteases and displayed an overall amino acid identity between 25 and 66% as compared to aspartyl proteases from other organisms. In addition to the five aspartyl proteases detected at the protein level, we have identified the coding sequences for seven other enzymes of this protease family by a similarity search in the genomic DNA of C. elegans which has recently been completely sequenced.

Amino Acid Sequence↗

Molecular cloning and functional expression of the first two specific insect myosuppressin receptors.

The Drosophila Genome Project database contains the sequences of two genes, CG8985 and CG13803, which are predicted to code for G protein-coupled receptors. We cloned the cDNAs corresponding to these genes and found that their gene structures had not been correctly annotated. We subsequently expressed the coding regions of the two corrected receptor genes in Chinese hamster ovary cells and found that each of them coded for a receptor that could be activated by low concentrations of Drosophila myosuppressin (EC50,4 x 10(-8) M). The insect myosuppressins are decapeptides that generally inhibit insect visceral muscles. Other tested Drosophila neuropeptides did not activate the two receptors. In addition to the two Drosophila myosuppressin receptors, we identified a sequence in the genomic database from the malaria mosquito Anopheles gambiae that also very likely codes for a myosuppressin receptor. To our knowledge, this paper is the first report on the molecular identification of specific insect myosuppressin receptors.

Amino Acid Sequence↗

Nucleotide sequence of cloned cDNAs encoding human preproparathyroid hormone.

We have cloned cDNA copies of human preproparathyroid hormone in Escherichia coli after insertion of double-stranded DNA into the Pst I site of plasmid pBR322 using the poly(dG) . poly(dC) homopolymer extension technique. Recombinant plasmids coding for preproparathyroid hormone were identified by filter hybridization assay using as a probe 32P-labeled bovine preproparathyroid cDNA. Nucleotide sequence analysis of five recombinant plasmids permitted the assignment of 74 nucleotides of the 5' noncoding region, the entire coding region of 345 nucleotides, and the entire 3' noncoding region of 348 nucleotides of the mRNA. The coding sequence predicts the previously unknown "pre" amino acid sequence and clarifies the hormone's amino acid sequence, which has been disrupted. The 5' noncoding region contains an AUG codon followed by a UGA stop codon before the authentic initiator codon. The 3' noncoding region is 120 nucleotides longer than in bovine preproparathyroid mRNA and contains two A-A-U-A-A-A sequences, potential signals for polyadenylation.

Base Sequence↗

Human somatostatin I: sequence of the cDNA.

RNA has been isolated from a human pancreatic somatostatinoma and used to prepare a cDNA library. After prescreening, clones containing somatostatin I sequences were identified by hybridization with an anglerfish somatostatin I-cloned cDNA probe. From the nucleotide sequence of two of these clones, we have deduced an essentially full-length mRNA sequence, including the preprosomatostatin coding region, 105 nucleotides from the 5' untranslated region and the complete 150-nucleotide 3' untranslated region. The coding region predicts a 116-amino acid precursor protein (Mr, 12.727) that contains somatostatin-14 and -28 at its COOH terminus. The predicted amino acid sequence of human somatostatin-28 is identical to that of somatostatin-28 isolated from the porcine and ovine species. A comparison of the amino acid sequences of human and anglerfish preprosomatostatin I indicated that the COOH-terminal region encoding somatostatin-14 and the adjacent 6 amino acids are highly conserved, whereas the remainder of the molecule, including the signal peptide region, is more divergent. However, many of the amino acid differences found in the pro region of the human and anglerfish proteins are conservative changes. This suggests that the propeptides have a similar secondary structure, which in turn may imply a biological function for this region of the molecule.

Amino Acid Sequence↗

Cystic fibrosis gene expression is not correlated with rectifying Cl- channels.

Cystic fibrosis (CF) involves a profound reduction of Cl- permeability in several exocrine tissues. A distinctive, outwardly rectifying, depolarization-induced Cl- channel (ORDIC channel) has been proposed to account for the Cl- conductance that is defective in CF. The recently identified CF gene is predicted to code for a 1480-amino acid integral membrane protein termed the CF transmembrane conductance regulator (CFTR). The CFTR shares sequence similarity with a superfamily of ATP-binding membrane transport proteins such as P-glycoprotein and STE6, but it also has features consistent with an ion channel function. It has been proposed that the CFTR might be an ORDIC channel. To determine if CFTR and ORDIC channel expression are correlated, we surveyed various cell lines for natural variation in CFTR and ORDIC channel expression. In four human epithelial cell lines (T84, CaCo2, PANC-1, and 9HTEo-/S) that encompass the full observed range of CFTR mRNA levels and ORDIC channel density we found no correlation.

Base Sequence↗

Cloning of a sodium channel alpha subunit from rabbit Schwann cells.

Overlapping cDNA clones spanning the entire coding region of a Na-channel alpha subunit were isolated from cultured Schwann cells from rabbits. The coding region predicts a polypeptide (Nas) of 1984 amino acids exhibiting several features characteristic of Na-channel alpha subunits isolated from other tissues. Sequence comparisons showed that the Nas alpha subunit resembles most the family of Na channels isolated from brain (approximately 80% amino acid identity) and is least similar (approximately 55% amino acid identity) to the atypical Na channel expressed in human heart and the partial rat cDNA, NaG. As for the brain II and III isoforms, two variants of Nas exist that appear to arise by alternative splicing. The results of reverse transcriptase-polymerase chain reaction experiments suggest that expression of Nas transcripts is restricted to cells in the peripheral and central nervous systems. Expression was detected in cultured Schwann cells, sciatic nerve, brain, and spinal cord but not in skeletal or cardiac muscle, liver, kidney, or lung.

Alternative Splicing↗

Molecular analysis of the ovine cystic fibrosis transmembrane conductance regulator gene.

There is a need for a large-animal model to investigate the etiology and biology of cystic fibrosis (CF) lung disease and to study potential therapies. The development and electrophysiology of the sheep airway have been shown to exhibit close functional parallels with the human airway, particularly with respect to the respiratory epithelium. We have cloned and sequenced the ovine cystic fibrosis transmembrane conductance regulator (CFTR) cDNA. It shows a high degree of conservation at the DNA coding and predicted polypeptide levels with human CFTR: at the nucleic acid level there is a 90% conservation (compared with 80% between human and mouse CFTR cDNA); at the polypeptide level, the degree of similarity is 95% (compared with 88% between human and mouse). Northern blot analysis and reverse transcription-PCR have shown that the patterns of expression of the ovine CFTR gene are very similar to those seen in humans. Further, the developmental expression of CFTR in the sheep is equivalent to that observed in humans. Thus, overall a CF sheep should show lung pathology similar to that of humans with CF.

Amino Acid Sequence↗

Structure-based assignment of the biochemical function of a hypothetical protein: a test case of structural genomics.

Many small bacterial, archaebacterial, and eukaryotic genomes have been sequenced, and the larger eukaryotic genomes are predicted to be completely sequenced within the next decade. In all genomes sequenced to date, a large portion of these organisms' predicted protein coding regions encode polypeptides of unknown biochemical, biophysical, and/or cellular functions. Three-dimensional structures of these proteins may suggest biochemical or biophysical functions. Here we report the crystal structure of one such protein, MJ0577, from a hyperthermophile, Methanococcus jannaschii, at 1.7-A resolution. The structure contains a bound ATP, suggesting MJ0577 is an ATPase or an ATP-mediated molecular switch, which we confirm by biochemical experiments. Furthermore, the structure reveals different ATP binding motifs that are shared among many homologous hypothetical proteins in this family. This result indicates that structure-based assignment of molecular function is a viable approach for the large-scale biochemical assignment of proteins and for discovering new motifs, a basic premise of structural genomics.

Adenosine Triphosphatases↗

A beta-1,3-N-acetylglucosaminyltransferase with poly-N-acetyllactosamine synthase activity is structurally related to beta-1,3-galactosyltransferases.

Human and mouse cDNAs encoding a new beta-1, 3-N-acetylglucosaminyltransferase (beta3GnT) have been isolated from fetal and newborn brain libraries. The human and mouse cDNAs included ORFs coding for predicted type II transmembrane polypeptides of 329 and 325 aa, respectively. The human and mouse beta3GnT homologues shared 90% similarity. The beta3GnT gene was widely expressed in human and mouse tissues, although differences in the transcript levels were visible, thus indicating possible tissue-specific regulation mechanisms. The beta3GnT enzyme showed a marked preference for Gal(beta1-4)Glc(NAc)-based acceptors, whereas no activity was detected on type 1 Gal(beta1-3)GlcNAc and O-glycan core 1 Gal(beta1-3)GalNAc acceptors. The new beta3GnT enzyme was capable of both initiating and elongating poly-N-acetyllactosamine chains, which demonstrated its identity with the poly-N-acetyllactosamine synthase enzyme (E.C. 2.4.1.149), showed no similarity with the i antigen beta3GnT enzyme described recently, and, strikingly, included several amino acid motifs in its protein that have been recently identified in beta-1,3-galactosyltransferase enzymes. The comparison between the new UDP-GlcNAc:betaGal beta3GnT and the three UDP-Gal:betaGlcNAc beta-1,3-galactosyltransferases-I, -II, and -III reveals glycosyltransferases that share conserved sequence motifs though exhibiting inverted donor and acceptor specificities. This suggests that the conserved amino acid motifs likely represent residues required for the catalysis of the glycosidic (beta1-3) linkage.

Amino Acid Sequence↗