Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “noncoding genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Noncoding plastid trnT-trnF sequences reveal a well resolved phylogeny of basal angiosperms.

Recent contributions from DNA sequences have revolutionized our concept of systematic relationships in angiosperms. However, parts of the angiosperm tree remain unclear. Previous studies have been based on coding or rDNA regions of relatively conserved genes. A phylogeny for basal angiosperms based on noncoding, fast-evolving sequences of the chloroplast genome region trnT-trnF is presented. The recognition of simple direct repeats allowed a robust alignment. Mutational hot spots appear to be confined to certain sectors, as in two stem-loop regions of the trnL intron secondary structure. Our highly resolved and well-supported phylogeny depicts the New Caledonian Amborella as the sister to all other angiosperms, followed by Nymphaeaceae and an Austrobaileya-Illicium-Schisandra clade. Ceratophyllum is substantiated as a close relative of monocots, as is a monophyletic eumagnoliid clade consisting of Piperales plus Winterales sister to Laurales plus Magnoliales. Possible reasons for the striking congruence between the trnT-trnF based phylogeny and phylogenies generated from combined multi-gene, multi-genome data are discussed.

Base Sequence↗

Complete nucleotide sequence of an attenuated hepatitis A virus: comparison with wild-type virus.

The complete nucleotide sequence of an attenuated hepatitis A virus, HAV HM-175/7 MK-5, was determined from cloned cDNA. This virus was derived from wild-type HAV HM-175 after 32 passages in African green monkey kidney cells. The resultant cell culture-adapted virus is attenuated for chimpanzees. This virus was passaged an additional three times in monkey kidney cells to obtain sufficient virus for molecular cloning and was designated HM-175/7 MK-5. Three overlapping cDNA clones were obtained that together spanned the entire genome. Comparison of the nucleotide sequence of cDNA from wild-type virus (propagated in marmoset liver in vivo) with attenuated virus (grown in cell culture) showed 24 nucleotide changes distributed throughout the genome. Five base deletions occurred in the 5' noncoding region, and 12 of the 16 base substitutions in the coding region resulted in amino acid changes. Amino acid changes occurred in viral capsid proteins VP1 and VP2 and several of the nonstructural proteins. Thus, a small number of nucleotide changes are responsible for adaptation to cell culture and attenuation of HAV strain HM-175.

Animals↗

Two versions of the gene encoding the 41-kilodalton subunit of the telomere binding protein of Oxytricha nova.

Macronuclear chromosomes of the ciliated protozoan Oxytricha nova terminate with a single-stranded (T4G4)2 overhang. The (T4G4)2 telomeric overhang is tenaciously bound by a protein heterodimer. We have cloned and sequenced the gene encoding the 41-kDa subunit of this telomere binding protein. The predicted amino acid sequence comprises two distinct regions, a carboxyl-terminal two-thirds that is 23% lysine and bears similarity to histone H1 and an amino-terminal one-third containing a hydrophobic stretch of about 15 amino acids. Two macronuclear versions of the gene differ in nucleotide sequence at several positions, but the derived polypeptides differ only at a single position, Ser-110 or Ala-110. Both versions harbor a small intron. The existence of this intron demonstrates that, despite the elimination of 95% of the micronuclear genome from the developing macronucleus, at least some noncoding DNA is retained during macronuclear development of hypotrichous ciliates.

Amino Acid Sequence↗

Island colonization and evolution of the insular woody habit in Echium L. (Boraginaceae).

Numerous island-inhabiting species of predominantly herbaceous angiosperm genera are woody shrubs or trees. Such "insular woodiness" is strongly manifested in the genus Echium, in which the continental species of circummediterranean distribution are herbaceous, whereas endemic species of islands along the Atlantic coast of north Africa are woody perennial shrubs. The history of 37 Echium species was traced with 70 kb of noncoding DNA determined from both chloroplast and nuclear genomes. In all, 239 polymorphic positions with 137 informative sites, in addition to 27 informative indels, were found. Island-dwelling Echium species are shown to descend from herbaceous continental ancestors via a single island colonization event that occurred < 20 million years ago. Founding colonization appears to have taken place on the Canary Islands, from which the Madeira and Cape Verde archipelagos were invaded. Colonization of island habitats correlates with a recent origin of perennial woodiness from herbaceous habit and was furthermore accompanied by intense speciation, which brought forth remarkable diversity of forms among contemporary island endemics. We argue that the origin of insular woodiness involved response to counter-selection of inbreeding depression in founding island colonies.

Adaptation, Physiological↗

Ribosomal protein S5 interacts with the internal ribosomal entry site of hepatitis C virus.

Translational initiation of hepatitis C virus (HCV) genome RNA occurs via its highly structured 5' noncoding region called the internal ribosome entry site (IRES). Recent studies indicate that HCV IRES and 40 S ribosomal subunit form a stable binary complex that is believed to be important for the subsequent assembly of the 48 S initiation complex. Ribosomal protein (rp) S9 has been suggested as the prime candidate protein for binding of the HCV IRES to the 40 S subunit. RpS9 has a molecular mass of approximately 25 kDa in UV cross-linking experiments. In the present study, we examined the approximately 25-kDa proteins of the 40 S ribosome that form complexes with the HCV IRES upon UV cross-linking. Immunoprecipitation with specific antibodies against two 25-kDa 40 S proteins, rpS5 and rpS9, clearly identified rpS5 as the protein bound to the IRES. Thus, our results support rpS5 as the critical element in positioning the HCV RNA on the 40 S ribosomal subunit during translation initiation.

5' Untranslated Regions↗

Algorithms for extracting structured motifs using a suffix tree with an application to promoter and regulatory site consensus identification.

This paper introduces two exact algorithms for extracting conserved structured motifs from a set of DNA sequences. Structured motifs may be described as an ordered collection of p > or = 1 "boxes" (each box corresponding to one part of the structured motif), p substitution rates (one for each box) and p - 1 intervals of distance (one for each pair of successive boxes in the collection). The contents of the boxes--that is, the motifs themselves--are unknown at the start of the algorithm. This is precisely what the algorithms are meant to find. A suffix tree is used for finding such motifs. The algorithms are efficient enough to be able to infer site consensi, such as, for instance, promoter sequences or regulatory sites, from a set of unaligned sequences corresponding to the noncoding regions upstream from all genes of a genome. In particular, both algorithms time complexity scales linearly with N2n where n is the average length of the sequences and N their number. An application to the identification of promoter and regulatory consensus sequences in bacterial genomes is shown.

Algorithms↗

Correlation of genetic and physical maps at the A mating-type locus of Coprinus cinereus.

The A mating type locus of Coprinus cinereus is remarkable for its extreme diversity, with over 100 different alleles in natural populations. Classical genetic studies have demonstrated that this hypervariability arises in part from recombination between two subloci of A, alpha and beta, although more recent population genetic data have indicated a third segregating sublocus. In this study, we characterized the molecular basis by which recombination generates nonparental A mating types. We mapped the frequency and location of all recombination events in two crosses and correlated the genetic and physical maps of A. We found that all recombination events were located in 6 kb of noncoding DNA between the alpha and beta subloci and that the rate of recombination in this noncoding region matched that generally observed for this genome. No recombination within gene clusters or within coding regions was observed, and the two alpha and beta subloci described in genetic analyses correlated with the previously characterized alpha and beta gene clusters. We propose that pairs of genes constitute both the sex determining and the hereditary unit of A.

Chromosome Mapping↗

Evolutionary conservation of regulatory elements in vertebrate Hox gene clusters.

Comparisons of DNA sequences among evolutionarily distantly related genomes permit identification of conserved functional regions in noncoding DNA. Hox genes are highly conserved in vertebrates, occur in clusters, and are uninterrupted by other genes. We aligned (PipMaker) the nucleotide sequences of the HoxA clusters of tilapia, pufferfish, striped bass, zebrafish, horn shark, human, and mouse, which are separated by approximately 500 million years of evolution. In support of our approach, several identified putative regulatory elements known to regulate the expression of Hox genes were recovered. The majority of the newly identified putative regulatory elements contain short fragments that are almost completely conserved and are identical to known binding sites for regulatory proteins (Transfac database). The regulatory intergenic regions located between the genes that are expressed most anteriorly in the embryo are longer and apparently more evolutionarily conserved than those at the other end of Hox clusters. Different presumed regulatory sequences are retained in either the Aalpha or Abeta duplicated Hox clusters in the fish lineages. This suggests that the conserved elements are involved in different gene regulatory networks and supports the duplication-deletion-complementation model of functional divergence of duplicated genes.

Animals↗

Generation and analysis of nonhomologous RNA-RNA recombinants in brome mosaic virus: sequence complementarities at crossover sites.

All three single-stranded RNAs of the brome mosaic virus (BMV) genome contain a highly conserved, 193-base 3' noncoding region. To study the recombination between individual BMV RNA components, barley plants were infected with a mixture of in vitro-transcribed wild-type BMV RNAs 1 and 2 and an RNA3 mutant that carried a deletion near the 3' end. This generated a population of both homologous and nonhomologous 3' recombinant BMV RNA3 variants. Sequencing revealed that these recombinants were derived by either single or double crossovers with BMV RNA1 or RNA2. The primary sequences at recombinant junctions did not show any similarity. However, they could be aligned to form double-stranded heteroduplexes. This suggested that local hybridizations among BMV RNAs may support intermolecular exchanges.

Base Sequence↗

Cell proteins bind specifically to West Nile virus minus-strand 3' stem-loop RNA.

The first 96 nucleotides of the 5'noncoding region (NCR) of West Nile virus (WNV) genomic RNA were previously reported to form thermodynamically predicted stem-loop (SL) structures that are conserved among flaviviruses. The complementary minus-strand 3' NCR RNA, which is thought to function as a promoter for the synthesis of plus-strand RNA, forms a corresponding predicted SL structure. RNase probing of the WNV 3' minus-strand stem-loop RNA [WNV (-)3' SL RNA] confirmed the existence of a terminal secondary structure. RNA-protein binding studies were performed with BHK S100 cytoplasmic extracts and in vitro-synthesized WNV (-)3' SL RNA as the probe. Three RNA-protein complexes (complexes 1,2, and 3) were detected by a gel mobility shift assay, and the specificity of the RNA-protein interactions was confirmed by gel mobility shift and UV-induced cross-linking competition assays. Four BHK cell proteins with molecular masses of 108, 60, 50, and 42 kDa were detected by UV-induced cross-linking to the WNV (-)3' SL RNA. A preliminary mapping study indicated that all four proteins bound to the first 75 nucleotides of the WNV 3' minus-strand RNA, the region that contains the terminal SL. A flavivirus resistance phenotype was previously shown to be inherited in mice as a single, autosomal dominant allele. The efficiencies of infection of resistant cells and susceptible cells are similar, but resistant cells (C3H/RV) produce less genomic RNA than congenic, susceptible cells (C3H/He). Three RNA-protein complexes and four UV-induced cross-linked cell proteins with mobilities identical to those detected in BHK cell extracts with the WNV (-)3' SL RNA were found in both C3H/RV and C3H/He cell extracts. However, the half-life of the C3H/RV complex 1 was three times longer than that of the C3H/He complex 1. It is possible that the increased binding activity of one of the resistant cell proteins for the flavivirus minus-strand RNA could result in a reduced synthesis of plus-strand RNA as observed with the flavivirus resistance phenotype.

Animals↗

Prevalence of GB virus C/hepatitis G virus ribonucleic acid and anti-hepatitis G virus-E2 antibodies among Japanese children with histories of transfusions or with liver diseases.

To clarify the prevalence of Japanese children thought to be at a risk for infection with GB virus-C (GBV-C)/hepatitis G virus (HGV), we investigated the detection rates of serum GBV-C/ HGV ribonucleic acid (RNA) by reverse transcription-seminested PCR and serum anti-HGV-E2 antibody by ELISA in 162 children with histories of blood or plasma product transfusions or with liver diseases and performed phylogenetic analysis of the 5' noncoding region sequences of GBV-C/HGV genomes. Children with histories of transfusions were divided into those who had been treated with antineoplastic agents for malignant diseases (malignant group) and those who had received transfusions for nonmalignant diseases (nonmalignant group). Children with liver diseases were divided into hepatitis B (HBV), hepatitis C (HCV), and non-A-C hepatitis groups. We detected GBV-C/ HGV RNA in 11 of 33 (33.3%) and anti-HGV-E2 in 1 of 27 (3.7%) children in the malignant group and in 3 of 56 (5.4%) and 1 of 53 (1.9%) children, respectively, in the nonmalignant group. Neither GBV-C/HGV RNA nor anti-HGV-E2 was detected in the HBV and non-A-C hepatitis groups. GBV-C/HGV RNA and anti-HGV-E2 were detected in 7 of 23 (30.4%) and in 1 of 18 (5.6%) children, respectively, in the HCV group. All children positive for either GBV-C/HGV RNA or anti-HGV-E2, except one whose route of GBV-C/HGV infection suggested mother-to-infant transmission, had histories of transfusions. The phylogenetic analysis showed that all isolates in this study were divisible into three groups and that most of them were clustered into group 3 (Asian group).

Adolescent↗

Detection of mitochondrial DNA mutations in primary breast cancer and fine-needle aspirates.

To determine the frequency and distribution of mitochondrial DNA mutations in breast cancer, 18 primary breast tumors were analyzed by direct sequencing. Twelve somatic mutations not present in matched lymphocytes and normal breast tissues were detected in 11 of the tumors screened (61%). Of these mutations, five (42%) were deletions or insertions in a homopolymeric C-stretch between nucleotides 303-315 (D310) within the D-loop. The remaining seven mutations (58%) were single-base substitutions in the coding (ND1, ND4, ND5, and cytochrome b genes) or noncoding regions (D-loop) of the mitochondrial genome. In three cases (25%), the mutations detected in coding regions led to amino acid substitutions in the protein sequence. We then screened an additional 46 primary breast tumors with a rapid PCR-based assay to identify poly-C alterations in D310, and we found seven more cancers with alterations. Using D310 mutations as clonal marker, we detected identical changes in five of five matched fine-needle aspirates and in four of four metastases-positive lymph nodes. The high frequency of D310 alterations in primary breast cancer combined with the high sensitivity of the PCR-based assays provides a new molecular tool for cancer detection.

Biopsy, Needle↗

The Tub alpha 3 gene from Zea mays: structure and expression in dividing plant tissues.

A gene (Tub alpha 3) coding for an alpha-Tub, expressed in dividing tissues, has been cloned from Zea mays. The deduced amino acid (aa) sequence, 450 aa long, is very similar to the other plant alpha-Tub (85-89% homology) so far reported, and in particular to the other two aa sequences (alpha 1-Tub and alpha 2-Tub) already published from the same species (93% homology). The genomic structure is also very similar, having three introns located at the same positions as in the Tub alpha 1 and Tub alpha 2 genes, one of them placed at the same position in the homologous genes from Arabidopsis thaliana. Nevertheless, the noncoding sequences are very different from the two other maize genomic sequences. In particular, no homology has been found either in the 5' upstream or in the 3'-untranslated sequences. Using specific 3' probes, it has been possible to detect the mRNA coded by this gene in many of the plant organs measured, but its highest abundance is observed in the organs rich in dividing cells, a pattern correlated with that of the histone H4-encoding gene. A cDNA clone has been identified in maize coleoptiles and sequenced, confirming the expression of the Tub alpha 3 in this organ. No preferential accumulation in any organ of the plant was found, in contrast with what was observed in the Tub alpha 1 and Tub alpha 2 genes already described. The Tub alpha gene family seems to consist in maize by at least two groups of homologous sequences, each one including a maximum of two or three coding units.

Amino Acid Sequence↗

The complete maternal and paternal mitochondrial genomes of the Mediterranean mussel Mytilus galloprovincialis: implications for the doubly uniparental inheritance mode of mtDNA.

The maternal (F) and paternal (M) mitochondrial genomes of the mussel Mytilus galloprovincialis have diverged by about 20% in nucleotide sequence but retained identical gene content and gene arrangement and similar nucleotide composition and codon usage bias. Both lack the ATPase8 subunit gene, have two tRNAs for methionine and a longer open-reading frame for cox3 than seen in other mollusks. Between the F and M genomes, tRNAs are most conserved followed by rRNAs and protein-coding genes, even though the degree of divergence varies considerably among the latter. Divergence at nad3 is exceptionally low most likely because this gene includes the origin of transcription of the lagging strand (O(L)). Noncoding regions are the least conserved with the notable exception of the central domain of the main control region and a segment of another noncoding region immediately following nad3. The amino acid divergence (14%) of the two genomes is smaller than in two other pairs of conspecific genomes that are available in GenBank, that of the clam Venerupis philippinarum (34%) and of the fresh water mussel Inversidens japanensis (50%), suggesting that doubly uniparental inheritance of mtDNA emerged at different times in the three species or that there has been a relatively recent replacement of the male genome by the female in the Mytilus line. The latter hypothesis is supported from phylogenetic and population studies of Mytilidae. That the M genome contains a full complement of genes with no premature termination codons argues against it being a selfish element that rides with the sperm. It is shorter than the F by 118 bp, which apparently cannot account for the postulated replicative advantage of this genome over the F in male gonads. The high similarity of the two genomes explains why the F genome may assume the role of the M genome, but it does not exclude the possibility that for this to happen some M-specific sequences must be transferred on to the F genome by means of recombination. If such sequences exist they would most likely be located in noncoding regions.

Animals↗

Identification of the 5' terminal sequence of the SAR-55 and MEX-14 strains of hepatitis E virus and confirmation that the genome is capped.

Hepatitis E virus (HEV) is a nonenveloped virus with a genome of single-stranded, positive-sense RNA. The 5' terminal sequence of two HEV strains (SAR-55 and MEX-14) was determined by a 5' RNA ligase-mediated rapid amplification of cDNA ends (RACE) method designed to select capped RNAs. The 5' noncoding region of the SAR-55 and MEX-14 strains were amplified, confirming that the genomic RNA of HEV is capped. The 5' noncoding region of the SAR-55 strain had 25 nucleotides, which is two less than reported for the Burmese strain, and that of the MEX-14 strain had 24 nucleotides, which is 21 more than reported previously [Huang et al., 1992].

5' Untranslated Regions↗

Mass spectrometry allows direct identification of proteins in large genomes.

Proteome projects seek to provide systematic functional analysis of the genes uncovered by genome sequencing initiatives. Mass spectrometric protein identification is a key requirement in these studies but to date, database searching tools rely on the availability of protein sequences derived from full length cDNA, expressed sequence tags or predicted open reading frames (ORFs) from genomic sequences. We demonstrate here that proteins can be identified directly in large genomic databases using peptide sequence tags obtained by tandem mass spectrometry. On the background of vast amounts of noncoding DNA sequence, identified peptides localize coding sequences (exons) in a confined region of the genome, which contains the cognate gene. The approach does not require prior information about putative ORFs as predicted by computerized gene finding algorithms. The method scales to the complete human genome and allows identification, mapping, cloning and assistance in gene prediction of any protein for which minimal mass spectrometric information can be obtained. Several novel proteins from Arabidopsis thaliana and human have been discovered in this way.

Amino Acid Sequence↗

Identification of Rare Noncoding Variants in Familial Nonmedullary Thyroid Carcinoma.

BACKGROUND: Familial nonmedullary thyroid carcinoma (FNMTC) occurs when three or more family members are affected by usually papillary thyroid carcinoma (PTC), the most common form of NMTC. While the heritability to NMTC is among the highest of all cancers, the genetic determinants among NMTC families are not well understood. Here, we aim to understand the contribution of rare noncoding germline variants in the etiology of FNMTC. METHODS: We previously reported whole-genome sequencing (WGS) and linkage analysis in 17 PTC families and reported on 41 protein-coding variants in 40 genes that cosegregated with PTC in 11 of the families. Herein, we further leveraged our WGS data to include noncoding variants in our analysis for all 17 families. We hypothesized that most of the pathogenic noncoding variants would be located in theoretical or empirically determined regulatory regions that demonstrate at a minimum, basal thyroid expression, a positive family linkage score, and co-segregation among PTC-affected individuals. To test this hypothesis, we adopted a unique filtering strategy to identify variants that occurred in known DNA elements and transcription factor binding sites, near regions known to impact on gene expression or splicing in thyroid tissue, and/or in characterized thyroid enhancers. We annotated variants using two analyses (ENCODE and transcription factor binding site) within the BasePlayer software. We separately analyzed (1) expression quantitative trait loci, (2) splicing quantitative trait loci, and (3) thyroid enhancers. We then ranked variants according to predicted pathogenicity and performed Sanger sequencing in all individuals of each family. RESULTS: In total, 121 variants were selected based on in-silico prediction and our custom ranking analysis in each pedigree. Of these, 56 variants showed cosegregation among all PTC-affected individuals and were absent from unaffected individuals. This included candidate variants from five of the six PTC families for whom no protein-coding variants were previously found. CONCLUSION: Our data suggest that noncoding variants are important in the etiology of FNMTC and provide a framework for identifying noncoding germline variants using a novel approach. Further studies are needed to functionally characterize these variants to better understand the molecular mechanism of their pathogenicity.

Humans↗

Comparing vertebrate whole-genome shotgun reads to the human genome.

Multi-species sequence comparisons are a very efficient way to reveal conserved genes. Because sequence finishing is expensive and time consuming, many genome sequences are likely to stay incomplete. A challenge is to use these fragmented data for understanding the human genome. Methods for using cross-species whole-genome shotgun sequence (WGS) for genome annotation are described in this paper. About one-half million high-quality rat WGS reads (covering 7.5% of the rat genome) generated at the Baylor College of Medicine Human Genome Sequencing Center were compared with the human genome. Using computer-generated random reads as a negative control, a set of parameters was determined for reliable interpretation of BLAST search results. About 10% of the rat reads contain regions that are conserved in the human genomic sequence and about one-third of these include known gene-coding regions. Mapping the conserved regions to human chromosomes showed a 23-fold enrichment for coding regions compared with noncoding regions. This approach can also be applied to other mammalian genomes for gene finding. These data predicted approximately 42,500 genes in the human, slightly more than reported previously.

Animals↗