Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

EV DNA from pancreatic cancer patient-derived cells harbors molecular, coding, non-coding signatures and mutational hotspots.

DNA packaged into cancer cell-derived EV is not well appreciated. Here, we uncovered signatures of EV DNA secreted by pancreatic cancer cells. The cancer cells and non-cancer counterparts exhibit distinct low vs. high molecular weight (LMW vs. HMW) EV DNA fragments distribution, respectively. Genome sequencing and Single Nucleotide Variants analysis revealed that 95% of reads and 94% of SNVs map to noncoding regions of the genome. Given that ~1% of the human genome represents coding regions, the 5% mapping rate to coding regions suggests a non-random enrichment of certain coding regions and mutations. The LMW DNA fragments not only set cancer cells apart, but also harbor cancer specific enrichment of unique coding regions, the top nine being FAM135B, COL22A1, TSNARE1, KCNK9, ZFAT, JRK, MROH5, GSDMD, and MIR3667HG. Additionally, the cancer cells' LMW DNA fragments exhibit dense centromeric mapping more strikingly on chromosomes 3, 7, 9, 10, 11, 13, 17, and 20. Mutational profiling turned up close to 200 mutations specific for the cancer cells. Altogether, our analyses suggest that centromeric regions might hold clues to EV DNA content from pancreatic cancer, the molecular, mutational signatures thereof, and rationalizes the need for a new approach to DNA biomarker research.

Humans↗

Compositional correlations in canine genome reflects similarity with human genes.

The base compositional correlations that hold among various coding and noncoding regions of the canine genome have been analysed. The distribution pattern of genes, on the basis of GC(3) composition, shows a wide range similar to that observed in human. However the occurrence of maximum number of genes was observed in the range of 65-75% of GC(3) composition. The correlation between the coding DNA sequences of canine with the different noncoding regions (introns and flanking regions) is found to be significant and in many cases the degree of correlation show similarity to human genome. We found that these correlations are not limited to the GC content alone, but is holding at the level of the frequency of individual bases as well. The present study suggests that canines ideally belong to the predicted 'general mammalian pattern' of genome composition along with human beings.

Animals↗

Function and evolution of a minimal plastid genome from a nonphotosynthetic parasitic plant.

Complete nucleotide sequencing shows that the plastid genome of Epifagus virginiana, a nonphotosynthetic parasitic flowering plant, lacks all genes for photosynthesis and chlororespiration found in chloroplast genomes of green plants. The 70,028-base-pair genome contains only 42 genes, at least 38 of which specify components of the gene-expression apparatus of the plastid. Moreover, all chloroplast-encoded RNA polymerase genes and many tRNA and ribosomal protein genes have been lost. Since the genome is functional, nuclear gene products must compensate for some gene losses by means of previously unsuspected import mechanisms that may operate in all plastids. At least one of the four unassigned protein genes in Epifagus plastid DNA must have a nongenetic and nonbioenergetic function and, thereby, serve as the reason for the maintenance of an active genome. Many small insertions in the Epifagus plastid genome create tandem duplications and presumably arose by slippage mispairing during DNA replication. The extensive reduction in genome size in Epifagus reflects an intensification of the same processes of length mutation that govern the amount of noncoding DNA in chloroplast genomes. Remarkably, this massive pruning occurred with a virtual absence of gene order change.

Chromosomes↗

The primary structures of two yeast enolase genes. Homology between the 5' noncoding flanking regions of yeast enolase and glyceraldehyde-3-phosphate dehydrogenase genes.

Segments of yeast genomic DNA containing two enolase structural genes have been isolated by subculture cloning procedures using a cDNA hybridization probe synthesized from purified yeast enolase mRNA. Based on restriction endonuclease and transcriptional maps of these two segments of yeast DNA, each hybrid plasmid contains a region of extensive nucleotide sequence homology which forms hybrids with the cDNA probe. The DNA sequences which flank this homologous region in the two hybrid plasmids are nonhomologous indicating that these sequences are nontandemly repeated in the yeast genome. The complete nucleotide sequence of the coding as well as the flanking noncoding regions of these genes has been determined. The amino acid sequence predicted from one reading frame of both structural genes is extremely similar to that determined for yeast enolase (Chin, C. C. Q., Brewer, J. M., Eckard, E., and Wold, F. (1981) J. Biol. Chem. 256, 1370-1376), confirming that these isolated structural genes encode yeast enolase. The nucleotide sequences of the coding regions of the genes are approximately 95% homologous, and neither gene contains an intervening sequence. Codon utilization in the enolase genes follows the same biased pattern previously described for two yeast glyceraldehyde-3-phosphate dehydrogenase structural genes (Holland, J. P., and Holland, M. J. (1980) J. Biol. Chem. 255, 2596-2605). DNA blotting analysis confirmed that the isolated segments of yeast DNA are colinear with yeast genomic DNA and that there are two nontandemly repeated enolase genes per haploid yeast genome. The noncoding portions of the two enolase genes adjacent to the initiation and termination codons are approximately 70% homologous and contain sequences thought to be involved in the synthesis and processing messenger RNA. Finally there are regions of extensive homology between the two enolase structural genes and two yeast glyceraldehyde-3-phosphate dehydrogenase structural genes within the 5- noncoding portions of these glycolytic genes.

Base Sequence↗

Attenuated Mengo virus: a new vector for live recombinant vaccines.

Several features make Mengo virus an excellent candidate for use as a vaccine vector. The virus has a wide host range, including rodents, pigs, monkeys, and most likely humans, and expresses its genome exclusively in the cytoplasm of the infected cell. Stable attenuated strains exist which are deleted for part of the 5' noncoding region of the genome. Here we report an attenuated Mengo virus recombinant, vLCMG4, that encodes an immunodominant cytotoxic T-lymphocyte epitope of the lymphocytic choriomeningitis virus (LCMV) nucleo-protein. vLCMG4 induced protective immunity against lethal LCMV infection after a single, low-dose immunization in BALB/c mice and elicited an LCMV-specific CD8+ cytotoxic T lymphocyte response. This demonstrates the potential of recombinant Mengo virus vaccines to confer protection against infectious diseases by the induction of cellular immune responses.

Amino Acid Sequence↗

Comparison of Pax1/9 locus reveals 500-Myr-old syntenic block and evolutionary conserved noncoding regions.

Identification of conserved genomic regions within and between different genomes is crucial when studying genome evolution. Here, we described regions of strong synteny conservation between vertebrate deuterostomes (tetrapods and teleosts) and invertebrate deuterostomes (amphioxus and sea urchin). The shared gene contents across phylogenetically distant species demonstrate that the conservation of the regions stemmed from an ancestral segment instead of a series of independent convergent events. Comparison of the syntenic regions allows us to postulate the primitive gene organization in the last common ancestor of deuterostomes and the evolutionary events that occurred to the 3 distinct lineages of sea urchin, amphioxus, and vertebrates after their separation. In addition, alignment of the syntenic regions led to the identification of 8 noncoding evolutionarily conserved regions shared between amphioxus and vertebrates. To our knowledge, this is the first report of conserved noncoding sequences shared by vertebrates and nonvertebrates. These noncoding sequences have high possibility of being elements that regulate neighboring genes. They are likely to be a factor in the maintenance of conserved synteny over long phylogenetic distance in different deuterostome lineages.

Amino Acid Sequence↗

Transcription of bxd noncoding RNAs promoted by trithorax represses Ubx in cis by transcriptional interference.

Much of the genome is transcribed into long noncoding RNAs (ncRNAs). Previous data suggested that bithoraxoid (bxd) ncRNAs of the Drosophila bithorax complex (BX-C) prevent silencing of Ultrabithorax (Ubx) and recruit activating proteins of the trithorax group (trxG) to their maintenance elements (MEs). We found that, surprisingly, Ubx and several bxd ncRNAs are expressed in nonoverlapping patterns in both embryos and imaginal discs, suggesting that transcription of these ncRNAs is associated with repression, not activation, of Ubx. Our data rule out siRNA or miRNA-based mechanisms for repression by bxd ncRNAs. Rather, ncRNA transcription itself, acting in cis, represses Ubx. The Trithorax complex TAC1 binds the Ubx coding region in nuclei expressing Ubx, and the bxd region in nuclei not expressing Ubx. We propose that TAC1 promotes the mosaic pattern of Ubx expression by facilitating transcriptional elongation of bxd ncRNAs, which represses Ubx transcription.

Animals↗

Analysis of genomic downsizing on the basis of region-of-difference polymorphism profiling of Mycobacterium tuberculosis patient isolates reveals geographic partitioning.

Mycobacterium tuberculosis, the etiological agent of tuberculosis, has lost many coding and noncoding regions in its genome during the course of evolution. We performed region-of-difference (RD) analysis using PCR-based genotyping of 131 M. tuberculosis clinical isolates obtained from four different countries, namely, India, Peru, Libya, and Angola. Our studies revealed that RD patterns are often distinct for strains circulating in specific geographical regions and can be used to trace the descent and spread of an isolate from its original reservoir. We describe our findings, which show that no single isolate from the four countries (n = 131) had all the 15 RDs either deleted or retained. Tuberculosis-specific deletion 1 (TbD1) was found to be conserved in 23% of the Indian isolates, indicating their possible ancient origin. RD9 was the most conserved region, RD11 was predominantly deleted, and RD6 was the most variable among the isolates in our collection irrespective of their geographic region. In contrast to earlier reports, our results demonstrate that the deletion of RD1 does not correlate with a decrease in the virulence potential of M. tuberculosis, as Indian isolates (n = 30) examined by us were from diseased individuals and yet had lost the RD1 region. Our results further illustrated that the intactness of the RD5 region may be associated with increased virulence of the organism. This study highlights that the RDs in M. tuberculosis genomes are geographically distributed and specific and may possibly be associated with virulence spectrum.

Angola↗

Genome-wide analysis of mammalian DNA segment fusion/fission.

As a powerful tool for gene function prediction, gene fusion has been widely studied in prokaryotes and certain groups of eukaryotes, but it has been little applied in studies of mammalian genomes. With the first fully sequenced mammalian genomes (human, mouse, rat) now available, we defined and collected a set of fusion/fission event-linked segments (FFLS) based on structured organized genomic alignment. The statistics of the sequence features highlighted the FFLSs against their random context. We found that there are three groups of FFLSs with different component pairs (i.e. gene-gene, gene-noncoding and noncoding-noncoding) in all three mammalian genomes. The proteins encoded by the components of FFLSs in the first group shown a strong tendency to interact with each other. The segmental components in the last two groups which did not contain any protein-coding genes, were found not only to be transcribed to some level, but also more conserved than the random background. Thus, these segments are possibly carrying certain biologically functional elements. We propose that FFLS may be a potential tool for prediction and analysis of function and functional interaction of genetic elements, including both genes and noncoding elements, in mammalian genomes. The full list of the FFLSs in the genomes of the three mammals is available as supporting information at doi:10.1016/j.jtbi.2005.09.016.

Animals↗

Sequence of the genome of lactate dehydrogenase-elevating virus: heterogenicity between strains P and C.

The complete nucleotide sequence of genomic RNA (14104 nt) of one strain of lactate dehydrogenase-elevating virus (LDV), LDV-P, is reported. It exhibits only about 80% nucleotide identity with the sequence reported for another LDV strain, LDV-C (Godeny et al., Virology 194, 585-596 (1993), and is 68 nucleotides shorter than the reported LDV-C sequence. The difference in length is largely due to the lack of a 59-nucleotide-long direct repeat in ORF 1a of the reported LDV-C sequence. Sequence analysis of a total of 1.4 kb of ORF 1a of LDV-C via reverse transcription/polymerase chain reaction (RT/PCR) technology failed to confirm the presence of this repeat in the LDV-C genome as well as of 24 deletions/insertions of single nucleotides that give rise to apparent transient reading frame differences between the LDV-P and LDV-C genomes and might have represented frameshift mutations. An additional 35 nucleotides in ORF 1a of the RT/PCR LDV-C products were the same as in the LDV-P rather than the reported LDV-C genome. The nucleotide sequences of the 5' leader and the 3' noncoding ends of the two genomes and the heptanucleotides involved in joining the 5' leader to the bodies of the subgenomic mRNAs were highly conserved or identical. The predicted LDV-P proteins, however, differed from those predicted for the LDV-C proteins between 25% for the ORF 2 protein and 1% for the ORF 7 nucleocapsid protein. All functional motifs of the ORF 1a and ORF 1b proteins were conserved. The ORF 1a protein possesses 11 potential transmembrane segments that flank the serine protease domain.

Amino Acid Sequence↗

Genetic analysis of the VP1 region of human enterovirus 71 strains isolated in Korea during 2000.

We have isolated Human enterovirus 71 (EV71) from stool and CSF samples taken from patients with acute flaccid paralysis, herpangina, or hand, foot and mouth disease in 2000. Both the cell culture-neutralization test and RT-PCR were used to detect enteroviruses. Rhabdomyosarcoma (RD), HEP2c, and BGM cells were used for the isolation of viruses, and serotypes were determined by the neutralization test using EV71-specific antiserum. For genomic analysis, we amplified a 437-bp fragment of the 5'-noncoding region of the enterovirus genome and a 484-bp fragment of the VP3/VP1 region of EV71 by RT-PCR, with positive results. Products amplified using an EV71-specific primer pair were sequenced and compared with other isolates of EV71. Analysis of the nucleotide sequences of the amplified fragments showed that the EV71 isolates from patients were over 98% homologous and belonged to the genotype C.

Amino Acid Sequence↗

Large-scale methylation patterns in the nuclear genomes of plants.

Methylation was investigated in compositional fractions of nuclear DNA preparations (50-100 kb in size) from five plants (onion, maize, rye, pea and tobacco), and was found to increase from GC-poor to GC-rich fractions. This methylation gradient showed different patterns in different plants and appears, therefore, to represent a novel, characteristic genome feature which concerns the noncoding, intergenic sequences that make up the bulk of the plant genomes investigated and mainly consist of repetitive sequences. The structural and functional implications of these results are discussed.

5-Methylcytosine↗

Relationship between alcoholic liver disease and HCV infection.

The high prevalence of hepatitis C virus (HCV) markers in alcoholic liver cirrhosis (AL-LC) and hepatocellular carcinoma (HCC) suggests a close aetiopathogenic relationship between alcoholic liver disease (ALD) and HCV infection. In the present study, HCV markers in ALD were measured by the highly sensitive methods, and the changes of sequential HCV markers after abstinence in ALD patients were analysed in order to elucidate the effect of alcohol on HCV. Antibodies to HCV-related antigen were determined using the first or second generation test kit. HCV-RNA genomes encoding the NS-5 region were detected using the RT-PCR method. In the HCV-NS5 negative serum, HCV genomes of the 5'-noncoding region were detected using the two-stage PCR method. Titres of HCV-RNA were measured by multiple cyclic PCR and cDNA dot blotting. Typing of HCV genomes was carried out on the PCR product from the NS-5 region by slot blot hybridization using type-specific cDNA probes, or by restriction fragment length polymorphisms analysis. In alcoholic fibrosis and alcoholic hepatitis, the prevalence of HCV markers was low, suggesting that the main aetiological factor is alcohol but not HCV in these types of ALD. HCV markers were positive in the half of the patients with AL-LC, and in more than 80% of patients with AL-CH and AL-HCC, indicating that HCV infection closely relates to these types of ALD. The ratio of the K1 type to the K2 type of HCV genomes was 4:1 in all types of NANB liver disease.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

The genome of the obligately intracellular bacterium Ehrlichia canis reveals themes of complex membrane structure and immune evasion strategies.

Ehrlichia canis, a small obligately intracellular, tick-transmitted, gram-negative, alpha-proteobacterium, is the primary etiologic agent of globally distributed canine monocytic ehrlichiosis. Complete genome sequencing revealed that the E. canis genome consists of a single circular chromosome of 1,315,030 bp predicted to encode 925 proteins, 40 stable RNA species, 17 putative pseudogenes, and a substantial proportion of noncoding sequence (27%). Interesting genome features include a large set of proteins with transmembrane helices and/or signal sequences and a unique serine-threonine bias associated with the potential for O glycosylation that was prominent in proteins associated with pathogen-host interactions. Furthermore, two paralogous protein families associated with immune evasion were identified, one of which contains poly(G-C) tracts, suggesting that they may play a role in phase variation and facilitation of persistent infections. Genes associated with pathogen-host interactions were identified, including a small group encoding proteins (n = 12) with tandem repeats and another group encoding proteins with eukaryote-like ankyrin domains (n = 7).

Animals↗

The infectious bronchitis virus nucleocapsid protein binds RNA sequences in the 3' terminus of the genome.

The infectious bronchitis virus (IBV) nucleocapsid protein was expressed as a bacterial fusion protein which differed from the native protein only in the addition of six amino terminus histidine residues. Using RNA overlay protein blot assays, the recombinant protein was shown to bind to RNA fragments specific for the positive sense 3' noncoding end of the IBV genome. At greater concentrations of sodium chloride, the native and fusion nucleocapsid proteins similarly bound to G RNA, representing the terminal 1805 3' nt of the genome, whereas bovine serum albumin and allantoic fluid protein did not bind to labeled G RNA. Competitive gel shift assays with labeled G RNA indicated that the protein interacted with several unlabeled RNA representing sequences at the 3' noncoding end of the IBV genome. Cache Valley virus (a bunyavirus) mRNA transcribed from the small segment cDNA also inhibited the interaction with IBV G RNA to approximately the same extent as homologous unlabeled G RNA, whereas reactions with bovine liver RNA and yeast tRNA were considerably weaker. Whereas yeast tRNA did not inhibit the interaction with the labeled large G RNA, interactions of the fusion protein with EF, a region from 78 to 217 nt from the 3' terminus of the IBV genome, were also apparently weaker than interactions with fragment CD which consisted of the 3' terminal 155 nt. On a molar basis, the latter interacted in an identical nature to a RNA consisting of CD and an additional 1053 nt of plasmid sequences. Compared to bovine liver RNA, unlabeled G specifically inhibited binding to the two smaller labeled IBV fragments in gel shift assays. The binding of IBV nucleocapsid protein with RNA probably requires specific sequences and/or structures that are present on the genome, and may represent a common mechanism used by similar viral nucleoproteins whose functions depend on binding to RNA.

Binding, Competitive↗

Infectious hypodermal and hematopoietic necrosis virus of shrimp is related to mosquito brevidensoviruses.

We purified and sequenced infectious hypodermal and hematopoietic necrosis virus (IHHNV), a small DNA virus of shrimp, from wild Penaeus stylirostris. The virion has a buoyant density of 1.45 as determined by cesium chloride gradient. Analysis of 3873 nucleotides of the viral genome revealed three large open reading frames (ORFs) and parts of the noncoding termini of the viral genome. The left, mid, and right ORFs on the complementary (plus) strand have potential coding capacities of 666 amino acids (aa) (75.77 kDa), 363 aa (42.11 kDa), and 329 aa (37.48 kDa), respectively. The overall genomic organization is similar to that of the mosquito brevidensoviruses. The left ORF most likely encodes the major nonstructural (NS) protein (NS-1) since it contains conserved replication initiator motifs and NTP-binding and helicase domains similar to those in NS-1 from all other parvoviruses. The IHHNV putative NS-1 shares the highest aa sequence homology with the NS-1 of mosquito brevidensoviruses, Aedes densovirus and Aedes albopictus parvovirus. A search for putative splicing sites revealed that the N-terminal region of NS-1 is very likely located in a small ORF upstream of the left ORF. The right ORF is presumed to encode structural polypeptides (VPs), as in other parvoviruses. Two putative promoters, located upstream of the left and right ORFs, are presumed to regulate expression of NS and VP genes, respectively. Thus, IHHNV is closely related to densoviruses of the genus Brevidensovirus in the family Parvoviridae, and we therefore propose to rename this virus Penaeus stylirostris densovirus (PstDNV).

Amino Acid Sequence↗

Molecular characterization of a cloned dolphin mitochondrial genome.

DNA clones have been isolated that span the complete mitochondrial (mt) genome of the dolphin, Cephalorhynchus commersonii. Hybridization experiments with purified primate mtDNA probes have established that there is close resemblance in the general organization of the dolphin mt genome and the terrestrial mammalian mt genomes. Sequences covering 2381 bp of the dolphin mt genome from the major noncoding region, three tRNA genes, and parts of the genes encoding cytochrome b, NADH dehydrogenase subunit 3 (ND3), and 16S rRNA have been compared with corresponding regions from other mammalian genomes. There is a general tendency throughout the sequenced regions for greater similarity between dolphin and bovine mt genomes than between dolphin and rodent or human mt genomes.

Animals↗

Rapid and accurate pyrosequencing of angiosperm plastid genomes.

BACKGROUND: Plastid genome sequence information is vital to several disciplines in plant biology, including phylogenetics and molecular biology. The past five years have witnessed a dramatic increase in the number of completely sequenced plastid genomes, fuelled largely by advances in conventional Sanger sequencing technology. Here we report a further significant reduction in time and cost for plastid genome sequencing through the successful use of a newly available pyrosequencing platform, the Genome Sequencer 20 (GS 20) System (454 Life Sciences Corporation), to rapidly and accurately sequence the whole plastid genomes of the basal eudicot angiosperms Nandina domestica (Berberidaceae) and Platanus occidentalis (Platanaceae). RESULTS: More than 99.75% of each plastid genome was simultaneously obtained during two GS 20 sequence runs, to an average depth of coverage of 24.6x in Nandina and 17.3x in Platanus. The Nandina and Platanus plastid genomes shared essentially identical gene complements and possessed the typical angiosperm plastid structure and gene arrangement. To assess the accuracy of the GS 20 sequence, over 45 kilobases of sequence were generated for each genome using conventional sequencing. Overall error rates of 0.043% and 0.031% were observed in GS 20 sequence for Nandina and Platanus, respectively. More than 97% of all observed errors were associated with homopolymer runs, with approximately 60% of all errors associated with homopolymer runs of 5 or more nucleotides and approximately 50% of all errors associated with regions of extensive homopolymer runs. No substitution errors were present in either genome. Error rates were generally higher in the single-copy and noncoding regions of both plastid genomes relative to the inverted repeat and coding regions. CONCLUSION: Highly accurate and essentially complete sequence information was obtained for the Nandina and Platanus plastid genomes using the GS 20 System. More importantly, the high accuracy observed in the GS 20 plastid genome sequence was generated for a significant reduction in time and cost over traditional shotgun-based genome sequencing techniques, although with approximately half the coverage of previously reported GS 20 de novo genome sequence. The GS 20 should be broadly applicable to angiosperm plastid genome sequencing, and therefore promises to expand the scale of plant genetic and phylogenetic research dramatically.

Base Sequence↗