Molecular diversity, structure and domestication of grasses.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The expression of type 1 fimbriae (pili) of Escherichia coli is turned on and off at the transcriptional level at a high frequency (10(-3) per cell per generation) in a process termed phase variation. Using Southern blot and DNA sequence analysis, we have detected a genomic rearrangement in the switch region immediately upstream of the fimbrial structural gene. This rearrangement involves an invertible 314-base-pair segment of DNA whose alternating orientation apparently results in the on-and-off activation of a promoter that determines the state of fimbrial expression.
It was shown recently that mutations of the ATRX gene give rise to a severe, X-linked form of syndromal mental retardation associated with alpha thalassaemia (ATR-X syndrome). In this study, we have characterised the full-length cDNA and predicted structure of the ATRX protein. Comparative analysis shows that it is an entirely new member of the SNF2 subgroup of a superfamily of proteins with similar ATPase and helicase domains. ATRX probably acts as a regulator of gene expression. Definition of its genomic structure enabled us to identify four novel splicing defects by screening 52 affected individuals. Correlation between these and previously identified mutations with variations in the ATR-X phenotype provides insights into the pathophysiology of this disease and the normal role of the ATRX protein in vivo.
We provide the first nucleotide sequence information for the slow isoform of troponin T (TnT). Sequence and hybridization analyses revealed that a single slow TnT gene present in the human genome gives rise to at least two different slow TnT variants by alternative splicing. The observed variations in slow TnT splicing generated major structural differences between the two corresponding slow TnT proteins in a domain that is likely to be involved in critical interactions with troponin C, troponin I, and tropomyosin in the thin filament. Corresponding variations have not been found for fast or for cardiac TnT. The comparison of splicing patterns for fast, cardiac, and slow TnT reveals that the splicing pattern for each isoform is unique. These features raise important questions of why and how all the individual members of the closely related TnT gene family developed such complex but different schemes of alternative splicing to create sets of variant proteins. This unusual familial trait is not known in any other muscle or nonmuscle multigene family.
Trypanosomes which parasitize mammals have evolved mechanisms to evade immune attack, such as the occupation of 'safe' intracellular sites (for example, Trypanosoma cruzi), or antigenic variation, exemplified by the salivarian trypanosomes (for example, Trypanosoma brucei). Antigenic variation is mediated by sequential expression of single variant surface glycoprotein (VSG) genes, and often involves transposition of the active gene. Every VSG transcript examined shares the same 5' terminal 35-nucleotide leader sequence. In T. brucei, this leader is encoded within a 1.4-kilobase unit tandemly reiterated to form a large array. It is hypothesized that this array is distantly linked to the expressed VSG gene and functions as a multiple promoter of VSG gene transcription, restricting transcription to that gene which, through genomic rearrangement, is placed downstream from the array. Leader and structural gene sequences are presumably juxtaposed by RNA splicing. Here we show that several trypanosomatids, both those which undergo antigenic variation (Trypanosoma congolense and Trypanosoma vivax) and those which do not (T. cruzi and Leptomonas collosoma), contain reiterated sequences homologous to the T. brucei spliced leader (SL). These results suggest that the SL, although utilized in VSG gene expression, is an ancestral sequence also used in the expression of other trypanosomatid genes.
The complement component C4 genes located in the major histocompatibility complex (MHC) class III region exhibit an unusually complex pattern of variations in gene number, gene size, and nucleotide polymorphism. Duplication or deletion of a C4 gene always concurs with its neighboring genes serine/threonine nuclear protein kinase RP, steroid 21-hydroxylase (CYP21), and tenascin (TNX), which together form a genetic unit termed the RCCX module. A detailed molecular genetic analysis of C4A and C4B and RCCX modular arrangements was correlated with immunochemical studies of C4A and C4B protein polymorphism in 150 normal Caucasians. The results show that bimodular RCCX has a frequency of 69%, whereas monomodular and trimodular RCCX structures account for 17.0 and 14.0%, respectively. Three quarters of C4 genes harbor the endogenous retrovirus HERV-K(C4). Partial deficiencies of C4A and C4B, primarily due to gene deletions and homoexpression of C4A proteins, have a combined frequency of 31.6%. This is probably the most common variation of gene dosage and gene size in human genomes. The seven RCCX physical variants create a great repertoire of haplotypes and diploid combinations, and a heterozygosity frequency of 69.4%. This phenomenon promotes the exchange of genetic information among RCCX constituents that is important in homogenizing the structural and functional diversities of C4A and C4B proteins. However, such length variants may cause unequal, interchromosomal crossovers leading to MHC-associated diseases. An analyses of the RCCX structures in 22 salt-losing, congenital adrenal hyperplasia patients revealed a significant increase in the monomodular structure with a long C4 gene linked to the pseudogene CYP21A, and bimodular structures with two CYP21A, which are likely generated by recombinations between heterozygous RCCX length variants.
Hemoglobin (Hb) genes have been cloned from three different species of ciliated protists, P. multimicronucleatum, P. triaurelia and P. jenningsi. Southern blotting of the genomic DNAs using the P. caudatum Hb cDNA showed both intraspecies variation in different stocks of P. caudatum and interspecies variation within the genus Paramecium. The isolated Hb genes were composed of 118, 117 and 117 codons, and interrupted by a short intron with 27, 29 and 29 bp at the same position, in P. multimicronucleatum, P. triaurelia and P. jenningsi, respectively. This suggests that the one-intron and two-exon structure has been conserved in the Hb genes in this genus. The amino acid sequences of the Paramecium Hbs were more than 87% identical to one another and homologous to those from the other ciliated protists Tetrahymena thermophila and T. pyriformis, the green alga Chlamydomonas eugametos, and the cyanobacterium Nostoc commune Hbs, all of which consist of about 120 amino acid residues (120-aa group). In particular, the amino acid sequences of the P. triaurelia and P. jenningsi Hbs were the same, although there were 20 nucleotide differences between the coding regions in the two genes. A maximum likelihood inference as to the phylogenetic relationships among these genes suggests that the Paramecium Hbs genes have evolved more rapidly than the other genes in the 120-aa group, and that P. triaurelia and P. genningsi are sibling species and the P. aurelia complex became a small cell after it separated from P. jenningsi.
Mitochondrial DNA (mt DNA) in cells of vertebrate organisms can assume an unusual triplex DNA structure known as the displacement loop (D loop). This triplex DNA structure forms when a partially replicated heavy strand of mtDNA (7S mtDNA) remains annealed to the light strand, displacing the native heavy strand in this region. The D-loop region contains the promoters for both heavy- and light-strand transcription as well as the origin of heavy-strand replication. However, the distribution of triplex and duplex forms of mtDNA in relation to respiratory activity of mammalian tissues has not been systematically characterized, and the functional significance of the D-loop structure is unknown. In comparisons of specialized muscle subtypes within the same species and of the same muscle subtype in different species, the relative proportion of D-loop versus duplex forms of mtDNA in striated muscle tissues of several mammalian species demonstrated marked variation, ranging from 1% in glycolytic fast skeletal fibers of the rabbit to 65% in the mouse heart. There was a consistent and direct correlation between the ratio of triplex to duplex forms of mtDNA and the capacity of these tissues for oxidative metabolism. The proportion of D-loop forms likewise correlated directly with mtDNA copy number, mtRNA abundance, and the specific activity of the mtDNA (gamma) polymerase. The D-loop form of mtDNA does not appear to be transcribed at greater efficiency than the duplex form, since the ratio of mtDNA copy number to mtRNA was unrelated to the proportion of triplex mtDNA genomes. However, tissues with a preponderance of D-loop forms tended to express greater levels of cytochrome b mRNA relative to mitochondrial rRNA transcripts, suggesting that the triplex structure may be associated with variations in partial versus full-length transcription of the heavy strand.
Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.
The distribution of structural alterations of the chloroplast genome found in grass chloroplast (cp) DNA in comparison with that of tobacco was systematically surveyed in the cpDNAs of monocots. Southern hybridization and/or PCR analyses for the detection of (1) three inversions in the large single-copy region, (2) loss of an intron in the rpoC1 gene, (3) an extra-sequence insertion in the rpoC2 gene, (4) the deletion of ORF2280, (5) rearrangements of the accD (ORF512) gene, and (6) non-reciprocal translocation of the rpl23 gene, were carried out on cpDNAs isolated from 58 species, 22 families, and 11 orders, which covered almost all families of monocots. These structural alterations of cpDNA mostly occurred at the family level. However, only part of the Restionaceae possessed the inversion that characterizes the lineage of grass differentiation. The order of mutational events made it possible to reconstruct grass phylogeny in monocots. Since no variations in structural alterations of the cpDNA were found among the Poaceae, grass plants were inferred to have originated from an ancestor harboring these structural alterations of the chloroplast genome. These phylogenetic relationships were supported by the sequence data of rbcL.
Expression arrays yield enormous amounts of data linking genes, via their cDNA sequences, to gene expression patterns. This now allows the characterisation of gene expression in normal and diseased tissues, as well as the response of tissues to the application of therapeutic reagents. Expression array data can be analysed with respect to the underlying protein sequences, which facilitates the precise determination of when and where certain groups of genes are expressed. More recent developments of clustering algorithms take additional parameters of the experimental set-up into account, focusing more directly on co-regulated set of genes. However, the information concerning transcriptional regulatory networks responsible for the observed expression patterns is not contained within the cDNA sequences used to generate the arrays. Regulation of expression is determined to a large extent by the promoter sequences of the individual genes (and/or enhancers). The complete sequence of the human genome now provides the molecular basis for the identification of many regulatory regions. Promoter sequences for specific cDNAs can be obtained reliably from genomic sequences by exon mapping. In the many cases in which cDNAs are 5'-incomplete, high quality promoter prediction tools can be used to locate promoters directly in the genomic sequence. Once sufficient numbers of promoter sequences have been obtained, a comparative promoter analysis of the co-regulated genes and groups of genes can be applied in order to generate models describing the higher order levels of transcription factor binding site organisation within these promoter regions. Such modules represent the molecular mechanisms through which regulatory networks influence gene expression, and candidates can be determined solely by bioinformatics. This approach also provides a powerful alternative for elucidating the functional features of genes with no detectable sequence similarity, by linking them to other genes on the basis of their common promoter structures.
We sequenced the first ca. 900 bp of the 5'-trnL(UAA)-trnV(UAC)/ndhJ region of the chloroplast DNA of different Microseris accessions in order to resolve homoplasious length variation detected in the trnL(UAA)-trnF(GAA) region. We found two to four tandemly repeated trnF genes in the species of Microseris (Asteraceae, Lactuceae) and two in their sister genus Uropappus. Sequences indicated nonhomologous transitions between two, three, and four trnF genes in different Microseris taxa. Independent origins of similar trnF copy numbers were inferred from a chloroplast phylogeny of Microseris. The taxa involved grow on separate continents, supporting parallel origins of similar length variants. The changes in trnF copy numbers were best explained by interchromosomal recombination with unequal crossing over. The 5' copies of the repeats showed the highest sequence conservation, suggesting that these copies are likely to be functional trnF genes, whereas the other ones probably represent pseudogenes. Our results show that length polymorphisms accumulate once a duplicated sequence has become incorporated. Due to parallel gains of similar trnF copy numbers, homoplasious length variation was introduced into the data matrix. The data demonstrate that length polymorphisms cannot be used as indicators for phylogenetic distance unless they can be analyzed at the sequence level.
The intron-genome size relationship was studied across a wide evolutionary range (from slime mold and yeast to human and maize), as well as the relationship between genome size and the ratio of intervening/coding sequence size. The average intron size is scaled to genome size with a slope of about one-fourth for the log-transformed values; i.e., on the global scale its increase in evolution is lower than the increase in genome size by four orders of magnitude. There are exceptions to the general trend. In baker's yeast introns are extraordinarily long for its genome size. Tetrapods also have longer introns than expected for their genome sizes. In teleost fish the mean intron size does not differ significantly, notwithstanding the differences in genome size. In contrast to previous reports, avian introns were not found to be significantly shorter than introns of mammals, although avian genomes are smaller than genomes of mammals on average by about a factor of 2.5. The extra-/intragenic ratio of noncoding DNA can be higher in fungi than in animals, notwithstanding the smaller fungal genomes. In vertebrates and invertebrates taken separately, this ratio is increasing as the increase in genome size. Two hypotheses are proposed to explain the variation in the extra-/intragenic ratio of noncoding DNA in organisms with similar numbers of genes: transition (dynamic) and equilibrium (static). According to the transition model, this variation arises with the rapid shift of genome size because the bulk of extragenic DNA can be changed more rapidly than the finely interspersed intron sequences. The equilibrium model assumes that this variation is a result of selective adjustment of genome size with constraints imposed on the intron size due to its putative link to chromatin structure (and constraints of the splicing machinery).
Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.
Comparison of the nucleotide sequences of the left arms of two Haemophilus influenzae phages, S2 and HP1 is presented. They exhibit a characteristic mosaic pattern of homologous and non-homologous regions. The homology extends over the attP site and int, orf 5 to 9, rep and the 3' part of cI genes. Two major non-homologous regions were detected. One is found between the int and cI genes; the other spans the region of promoters and the cox gene. Variations in the region of the promotors which is involved in the choice between a lysogenic and a lytic pathway and some divergences in the cI coding sequences are probably responsible for the observed immunity differences between the two phages. Distinctions in the distribution of consensus sequences for an integration host factor (IHF) and integrase-binding sites and promoters are described. These data offer an explanation of the relationship between three types of S2/HP1 phages. It allows in turn a final settlement of the nomenclature variation in the literature. The results presented, which are similar to those obtained for other phage groups, suggest that the mosaic structure of phage genomes is a normal outcome of phage divergence.
The mitochondrial genome (mtDNA) of the plant parasitic nematode Globodera pallida exists as a population of small, circular DNAs that, taken individually, are of insufficient length to encode the typical metazoan mitochondrial gene complement. As far as we are aware, this unusual structural organization is unique among higher metazoans, although interesting comparisons can be made with the multipartite mitochondrial genome organizations of plants and fungi. The variation in frequency between populations displayed by some components of the mtDNA is likely to have major implications for the way in which mtDNA can be used in population and evolutionary genetic studies of G. pallida.
Infectious diseases have shaped the human population genetic structure, and genetic variation influences the susceptibility to many viral diseases. However, a variety of challenges have made the implementation of traditional human Genome-wide Association Studies (GWAS) approaches to study these infectious outcomes challenging. In contrast, mouse models of infectious diseases provide an experimental control and precision, which facilitates analyses and mechanistic studies of the role of genetic variation on infection. Here we use a genetic mapping cross between two distinct Collaborative Cross mouse strains with respect to severe acute respiratory syndrome coronavirus (SARS-CoV) disease outcomes. We find several loci control differential disease outcome for a variety of traits in the context of SARS-CoV infection. Importantly, we identify a locus on mouse chromosome 9 that shows conserved synteny with a human GWAS locus for SARS-CoV-2 severe disease. We follow-up and confirm a role for this locus, and identify two candidate genes, CCR9 and CXCR6, that both play a key role in regulating the severity of SARS-CoV, SARS-CoV-2, and a distantly related bat sarbecovirus disease outcomes. As such we provide a template for using experimental mouse crosses to identify and characterize multitrait loci that regulate pathogenic infectious outcomes across species. IMPORTANCE Host genetic variation is an important determinant that predicts disease outcomes following infection. In the setting of highly pathogenic coronavirus infections genetic determinants underlying host susceptibility and mortality remain unclear. To elucidate the role of host genetic variation on sarbecovirus pathogenesis and disease outcomes, we utilized the Collaborative Cross (CC) mouse genetic reference population as a model to identify susceptibility alleles to SARS-CoV and SARS-CoV-2 infections. Our findings reveal that a multitrait loci found in chromosome 9 is an important regulator of sarbecovirus pathogenesis in mice. Within this locus, we identified and validated CCR9 and CXCR6 as important regulators of host disease outcomes. Specifically, both CCR9 and CXCR6 are protective against severe SARS-CoV, SARS-CoV-2, and SARS-related HKU3 virus disease in mice. This chromosome 9 multitrait locus may be important to help identify genes that regulate coronavirus disease outcomes in humans.
BACKGROUND/AIM: Colorectal cancer (CRC) remains a leading cause of cancer-related morbidity and mortality worldwide. Although immunotherapy has improved outcomes for a subset of patients, its limited efficacy in many cases highlights the need for a more comprehensive understanding of the CRC immune microenvironment. This study aimed to characterize the molecular landscape of the CRC immune microenvironment using an integrated multi-omics approach and to identify candidate regulatory molecules associated with immune remodelling. MATERIALS AND METHODS: We integrated structural variation, DNA methylation, chromatin accessibility, proteomic, and phosphoproteomic data generated from an in-house CRC cohort with transcriptomic data from The Cancer Genome Atlas (TCGA). Analyses focused on 1,539 immune-related genes (IRGs) associated with CD4+ T cells, B cells, and natural killer (NK) cells. Multi-layered genomic and proteomic analyses were performed to identify altered immune-related pathways, hub genes, candidate transcription factors, and upstream kinases. RESULTS: Higher infiltration of CD4+ T cells, B cells, and NK cells was associated with CRC. IRGs exhibited widespread alterations across genomic, epigenomic, transcriptomic, proteomic, and phosphoproteomic levels. IL10, LEP, ITGAM, and EGFR emerged as candidate hub genes. EGFR phosphorylation at S991 and T693 was significantly decreased in CRC. STAT2 and HSF1 were identified as candidate upstream transcription factors, while CDK2 emerged as a candidate upstream kinase associated with immune infiltration and immune checkpoint expression. CONCLUSION: This study provides a systematic multi-omics characterization of immune microenvironment remodelling in CRC and identifies candidate molecular regulators that may serve as potential targets for future immunotherapy research.