Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Structure and chromosome localization of the human eosinophil-derived neurotoxin and eosinophil cationic protein genes: evidence for intronless coding sequences in the ribonuclease gene superfamily.

Human genomic DNAs for the eosinophil granule proteins, eosinophil-derived neurotoxin (EDN) and eosinophil cationic protein (ECP), were isolated from genomic libraries. Alignment of EDN (RNS2) and ECP (RNS3) gene sequences demonstrated remarkable nucleotide similarities in noncoding sequences, introns, and flanking regions, as well as in the previously known coding regions. Detailed examination of the 5'-noncoding regions yielded putative TATA and CAAT boxes, as well as similarities to promoter motifs from unrelated genes. A single intron of 230 bases was found in the 5' untranslated region and we suggest that a single intron in this region and an intronless coding region are features common to many members of the RNase gene superfamily. The RNS2 and RNS3 genes were localized to the q24-q31 region of human chromosome 14. It is likely that these two genes arose as a consequence of a gene duplication event that took place approximately 25-40 million years ago and that a subset of anthropoid primates possess both of these genes or closely related genes.

Animals↗

Circadian clock control of ribosome composition promotes rhythmic translation and termination fidelity.

Ribosome composition is dynamic, shifting with cell state and stress, but whether it varies with circadian time is unknown. Here, we uncover circadian clock-driven changes in ribosome composition in Neurospora crassa. Mass spectrometry of ribosomes across circadian time identified six ribosomal proteins and one associated factor under clock control. Rhythms in eL31 abundance were validated in purified ribosomes, and deletion of el31 disrupted translation rhythms in nearly half of rhythmically translated mRNAs. N. crassa eL31 promotes circadian control of translation termination and impacts elongation fidelity while maintaining Mg homeostasis, a key determinant of translational accuracy. These findings reveal that the circadian clock reprograms ribosome composition to orchestrate rhythmic translation and fidelity, temporally expanding the proteome beyond the static genome to align cellular function with time of day.

Neurospora crassa↗

Transcriptional analysis of the acid-inducible asr gene in enterobacteria.

We show here that transcription of the asr gene in Escherichia coli, Salmonella enterica serovar Typhimurium, Klebsiella pneumoniae and Enterobacter cloacae is strongly dependent on the acidification level of the growth medium, with maximal induction at pH 4.0-4.5 as determined by Northern hybridization analysis. Previous gene array analyses have also shown that asr is the most acid-induced gene in the E. coli genome. Sequence alignment of the asr promoters from different enterobacterial species identified a highly conserved region located at position -70 to -30 relative to the asr transcriptional start site. By deletion of various segments of this region in the E. coli asr promoter it was shown that sequences upstream from the -40 position were important for induction. Transcription from the E. coli asr promoter was demonstrated to be growth-phase-dependent and to require the alternative sigma factor RpoS (sigma(S)) in stationary phase. Transcription of the asr gene was also found to be subject to negative control by the nucleoid protein H-NS.

Acids↗

De novo Genes in Plants: Origins, Mechanisms, and Functional Implications.

De novo genes originate from previously non-coding genomic regions. They provide an important source of lineage-specific innovation. In plants, these genes may contribute to adaptation, trait diversity and crop evolution. This review summarizes recent progress in plant de novo gene research. It first discusses major routes of gene birth, including transcription-first, open reading frame (ORF)-first and concurrent models. It also examines how nascent loci acquire regulatory control and enter existing biological networks. The review then summarizes their evolutionary features, including weak early constraint, rapid molecular change, restricted expression and structural refinement. It further discusses plant de novo genes involved in stress responses, seed germination, kernel dehydration, subspecies divergence, reproductive isolation and floral scent diversification. Current methods for identifying de novo genes remain limited by rapid sequence evolution, genome annotation quality, polyploidy and transposable elements. Whole-genome synteny alignment, multi-omics evidence and machine-learning approaches can improve candidate discovery. However, each method has important limitations. Finally, this review highlights key future questions in functional validation, latent coding potential in long non-coding RNAs, epigenetic activation, regulatory-network integration and crop improvement. These perspectives clarify how de novo genes shape plant adaptation and how they may be used in precision breeding and synthetic biology.

adaptive evolution↗

Phylogenetic footprinting reveals unexpected complexity in trans factor binding upstream from the epsilon-globin gene.

The human epsilon-globin gene undergoes dramatic changes in transcriptional activity during development, but the molecular factors that control its high expression in the embryo and its complete repression at 6-8 weeks of gestation are unknown. Although a putative silencer has been identified, the action of this silencer appears to be necessary but not sufficient for complete repression of epsilon gene expression, suggesting that multiple control elements may be required. Phylogenetic footprinting is a strategy that uses evolution to aid in the elucidation of these multiple control points. The strategy is based on the observation that the characteristic developmental expression pattern of the epsilon gene is conserved in all placental mammals. By aligning epsilon genomic sequences (from -2.0 kb upstream to the epsilon polyadenylylation signal), conserved sequence elements that are likely binding sites for trans factors can be identified against the background of neutral DNA. Twenty-one such conserved elements (phylogenetic footprints) were found upstream of the epsilon gene. Oligonucleotides spanning these conserved elements were used in a gel-shift assay to reveal 47 nuclear binding sites. Among these were 8 binding sites for YY1 (yin and yang 1), a protein with dual (activator or repressor) activity; 5 binding sites for the putative stage selector protein, SSP; and 7 binding sites for an as yet unidentified protein. The large number of high-affinity interactions detected in this analysis further supports the notion that the epsilon gene is regulated by multiple redundant elements.

Animals↗

A low rate of simultaneous double-nucleotide mutations in primates.

The occurrence of double-nucleotide (doublet) mutations is contrary to the normal assumption that point mutations affect single nucleotides. Here we develop a new method for estimating the doublet mutation rate and apply it to more than a megabase of human-chimpanzee-baboon genomic DNA alignments and more than a million human single-nucleotide polymorphisms. The new method accounts for the effect of regional variation in evolutionary rates, which may be a confounding factor in previous estimates of the doublet mutation rate. Furthermore we determine sequence context effects by using sequence comparisons over a variety of lineage lengths. This approach yields a new estimate of the doublet mutation rate of 0.3% of the singleton rate, indicating that doublet mutations are far rarer than previously thought. Our results suggest that doublet mutations are unlikely to have caused the correlation between synonymous and nonsynonymous substitution rates in mammals, and also show that regional variation and sequence context effects play an important role in primate DNA sequence evolution.

Animals↗

Melanoma-Reactive CD8+ T cells recognize a novel tumor antigen expressed in a wide variety of tumor types.

An autologous melanoma cell line selected for loss of expression of the immunodominant MART-1 and gp100 antigens was initially used to carry out a mixed lymphocyte tumor culture (MLTC) in a patient who expressed the human leukocyte antigen (HLA)-AI and HLA-A2 class I major histocompatibility complex alleles. Ten clones identified from this MLTC seemed to recognize melanoma in an HLA-A1-restricted manner but failed to recognize a panel of previously described melanoma antigens. The screening of an autologous melanoma cDNA library with one HLA-Al-restricted melanoma-reactive T-cell clone resulted in the isolation of a cDNA clone called AIM-2 (antigen isolated from immunoselected melanoma-2). The AIM-2 transcript seemed to have retained an intronic sequence based on its alignment with genomic sequences as well as expressed sequence tags. This transcript was not readily detected after Northern blot analysis of melanoma mRNA, indicating that only low levels of this product may be expressed in tumor cells. Quantitative reverse transcriptase-polymerase chain reaction analysis, however, demonstrated a correlation between T-cell recognition and expression in HLA-A1-expressing tumor cell lines. A peptide that was encoded within a short open reading frame of 23 amino acids and conformed to the HLA-A1 binding motif RSDSGQQARY was found to represent the T-cell epitope. The AIM-2-reactive T-cell clone recognized a number of neuroectodermal tumors as well as breast, ovarian, and colon carcinomas that expressed HLA-A1, indicating that this represents a widely expressed tumor antigen. Thus, AIM-2 may represent a potential target for the development of vaccines in patients bearing tumors of a variety of histologies.

Amino Acid Sequence↗

TOPAAS, a tomato and potato assembly assistance system for selection and finishing of bacterial artificial chromosomes.

We have developed the software package Tomato and Potato Assembly Assistance System (TOPAAS), which automates the assembly and scaffolding of contig sequences for low-coverage sequencing projects. The order of contigs predicted by TOPAAS is based on read pair information; alignments between genomic, expressed sequence tags, and bacterial artificial chromosome (BAC) end sequences; and annotated genes. The contig scaffold is used by TOPAAS for automated design of nonredundant sequence gap-flanking PCR primers. We show that TOPAAS builds reliable scaffolds for tomato (Solanum lycopersicum) and potato (Solanum tuberosum) BAC contigs that were assembled from shotgun sequences covering the target at 6- to 8-fold coverage. More than 90% of the gaps are closed by sequence PCR, based on the predicted ordering information. TOPAAS also assists the selection of large genomic insert clones from BAC libraries for walking. For this, tomato BACs are screened by automated BLAST analysis and in parallel, high-density nonselective amplified fragment length polymorphism fingerprinting is used for constructing a high-resolution BAC physical map. BLAST and amplified fragment length polymorphism analysis are then used together to determine the precise overlap. Assembly onto the seed BAC consensus confirms the BACs are properly selected for having an extremely short overlap and largest extending insert. This method will be particularly applicable where related or syntenic genomes are sequenced, as shown here for the Solanaceae, and potentially useful for the monocots Brassicaceae and Leguminosea.

Chromosomes, Artificial, Bacterial↗

Extensive sequence homology between the mycobacterium leprae LSR (12 kDa) antigen and its Mycobacterium tuberculosis counterpart.

The Mycobacterium leprae LSR (12 kDa) protein antigen has been reported to mimic whole cell M. leprae in T cell responses across the leprosy spectrum. In addition, B cell responses to specific sequences within the LSR antigen have been shown to be associated with immunopathological responses in leprosy patients with erythema nodosum leprosum. We have in the present study applied the M. leprae LSR DNA sequence as query to search for the presence of homologous genes within the recently completed Mycobacterium tuberculosis genome database (Sanger Centre, UK). By using the BLASTN search tool, a homologous M. tuberculosis open reading frame (336 bp), encoding a protein antigen of 12.1 kDa, was identified within the cosmid MTCY07H7B.25. The gene is designated Rv3597c within the M. tuberculosis H37Rv genome. Sequence alignment revealed 93% identity between the M. leprae and M. tuberculosis antigens at the amino acid sequence level. The finding that some B and T cell epitopes were localized to regions with amino acid substitutions may account for the putative differential responsiveness to this antigen in tuberculosis and leprosy.

Amino Acid Sequence↗

Functional and structural characterization of thermostable D-amino acid aminotransferases from Geobacillus spp.

D-amino acid aminotransferases (D-AATs) from Geobacillus toebii SK1 and Geobacillus sp. strain KLS1 were cloned and characterized from a genetic, catalytic, and structural aspect. Although the enzymes were highly thermostable, their catalytic capability was approximately one-third of that of highly active Bacilli enzymes, with respective turnover rates of 47 and 55 s(-1) at 50 degrees C. The Geobacillus enzymes were unique and shared limited sequence identities of below 45% with D-AATs from mesophilic and thermophilic Bacillus spp., except for a hypothetical protein with a 72% identity from the G. kaustophilus genome. Structural alignments showed that most key residues were conserved in the Geobacillus enzymes, although the conservative residues just before the catalytic lysine were distinctively changed: the 140-LRcD-143 sequence in Bacillus D-AATs was 144-EYcY-147 in the Geobacillus D-AATs. When the EYcY sequence from the SK1 enzyme was mutated into LRcD, a 68% increase in catalytic activity was observed, while the binding affinity toward alpha-ketoglutarate decreased by half. The mutant was very close to the wild-type in thermal stability, indicating that the mutations did not disturb the overall structure of the enzyme. Homology modeling also suggested that the two tyrosine residues in the EYcY sequence from the Geobacillus D-AATs had a pi/pi interaction that was replaceable with the salt bridge interaction between the arginine and aspartate residues in the LRcD sequence.

Amino Acid Sequence↗

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans↗

Characterization of mouse Atp6i gene, the gene promoter, and the gene expression.

Solubilization of bone mineral by osteoclasts depends on the formation of an acidic extracellular compartment through the action of a V-type ATPase. We previously cloned a gene encoding a putative osteoclast-specific proton pump subunit, termed OC-116 kDa, approved mouse Atp6i (ATPase, H+ transporting, [vacuolar proton pump] member I). The function of Atp6i as osteoclast-specific proton pump subunit was confirmed in our mouse knockout study. However, the transcription regulation of Atp6i remains largely unknown. In this study, the gene encoding mouse Atp6i and the promoter have been isolated and completely sequenced. In addition, the temporal and spatial expressions of Atp6i have been characterized. Intrachromosomal mapping studies revealed that the gene contains 20 exons and 19 introns spanning approximately 11 kilobases (kb) of genomic DNA. Alignment of the mouse Atp6i gene exon sequence and predicted amino acid sequence to that of the human reveals a strong homology at both the nucleotide (82%) and the amino acid (80%) levels. Primer extension assay indicates that there is one transcription start site at 48 base pairs (bp) upstream of the initiator Met codon. Analysis of 4 kb of the putative promoter region indicates that this gene lacks canonical TATA and CAAT boxes and contains multiple putative transcription regulatory elements. Northern blot analysis of RNAs from a number of mouse tissues reveals that Atp6i is expressed predominantly in osteoclasts, and this predominant expression was confirmed by reverse-transcription polymerase chain reaction (RT-PCR) assay and immunohistochemical analysis. Whole-mount in situ hybridization shows that Atp6i expression is detected initially in the headfold region and posterior region in the somite stage of mouse embryonic development (E8.5) and becomes progressively restricted to anterior regions and the limb bud by E9.5. The expression level of Atp6i is largely reduced after E10.5. This is the first report of the characterizations of Atp6i gene, its promoter, and its gene expression patterns during mouse development. This study may provide valuable insights into the function of Atp6i, its osteoclast-selective expression, regulation, and the molecular mechanisms responsible for osteoclast activation.

Adenosine Triphosphatases↗

Single-cell sequencing reveals synovial fluid γδ T-cell expansion in equine experimental osteoarthritis.

OBJECTIVE: Define temporal cellular changes following joint injury using single-cell RNA sequencing in experimental equine posttraumatic osteoarthritis (PTOA). METHODS: PTOA was induced in 4 Quarter Horses (3 to 5 years) via carpal osteochondral fragmentation and high-speed treadmill exercise. Synovial fluid (SF) cells and synovium were sampled over 18 weeks (November 2023 to April 2024). Single-cell suspensions were processed (10x Genomics Chromium iX), then aligned to the equine genome (Cell Ranger). Downstream analysis was completed in the R Seurat package. Differential gene expression (log2[fold change] > 1; P < .05) and differential abundance analyses were performed (P < .1). RESULTS: Cartilage injury had a modest impact on gene expression changes and cell abundance shifts in SF. Integrated analysis of 90,323 SF cells across 4 time points revealed 9 distinct cell types, primarily T cells (73 &#xb1; 19%) followed by myeloid cells (20 &#xb1; 13%). Subcluster analysis of T cells revealed 9 transcriptomically distinct subtypes (3 CD8, 2 CD4, 3 &#x3b3;&#x3b4;, and 1 cycling). Differential abundance analyses of temporal changes identified increased &#x3b3;&#x3b4; T and decreased CD4+ T-cell subsets in joints over time. Expanded populations of IL-23 receptor-positive &#x3b3;&#x3b4; T cells exhibited increased T-helper 17 signatures. CONCLUSIONS: IL-23 receptor-positive &#x3b3;&#x3b4; T-cell expansion, associated with joint inflammation, occurred in PTOA. Limitations include small sample size and individual heterogeneity; further investigation over extended timeframe is necessary to confirm whether later stages of the experimental model reflect natural chronic OA. CLINICAL RELEVANCE: Cellular immunotherapy targeting &#x3b3;&#x3b4; T cells and IL-23/IL-17 blockade may warrant investigation to mitigate equine OA progression.

equine↗

Dendritic cell-associated lectin-1: a novel dendritic cell-associated, C-type lectin-like molecule enhances T cell secretion of IL-4.

We have characterized dendritic cell (DC)-associated lectin-1 (DCAL-1), a novel, type II, transmembrane, C-type lectin-like protein. DCAL-1 has restricted expression in hemopoietic cells, in particular, DCs and B cells, but T cells and monocytes do not express it. The DCAL-1 locus is within a cluster of C-type lectin-like loci on human chromosome 12p12-13 just 3' to the CD69 locus. The consensus sequence of the DCAL-1 gene was confirmed by RACE-PCR; however, based on sequence alignment with genomic DNA and with various human expressed sequence tags, we predict that DCAL-1 has two splice variants. C-type lectins share a common sequence motif of 14 invariable and 18 highly conserved aa residues known as the carbohydrate recognition domain. DCAL-1, however, is missing three of the cysteine residues required to form the standard carbohydrate recognition domain. DCAL-1 mRNA and protein expression are increased upon the differentiation of monocytes to CD1a(+) DCs. B cells also express high levels of DCAL-1 on their cell surface. Using a DCAL-1 fusion protein we identified a population of CD4(+) CD45RA(+) T cells that express DCAL-1 ligand. Coincubation with soluble DCAL-1 enhanced the proliferation of CD4(+) T cells in response to CD3 ligation and significantly increased IL-4 secretion. In contrast, coincubation with soluble DC-specific ICAM-3-grabbing nonintegrin (CD209) fusion protein as a control had no effect on CD4(+) T cell proliferation or IL-4 and IFN-gamma secretion. Therefore, the function of DCAL-1 on DCs and B cells may act as a T cell costimulatory molecule, which skews CD4(+) T cells toward a Th2 response by enhancing their secretion of IL-4.

Adjuvants, Immunologic↗

Chromosome bands, their chromatin flavors, and their functional features.

To show that the input pattern of chromosomal mutations is highly organized relative to the band patterns along human chromosomes, a new term, "metaphase chromatin flavor," is introduced. Five different flavors of euchromatic metaphase bands are cytologically identified along a human ideogram. These are G-bands and, based upon combinations of extreme Alu richness and GC richness, four different R-band flavors. The two flavors with extremely GC-rich components, traditionally called "T-bands," represent only 15% of all bands. However, they contain 65% of mapped genes, 19 of 25 mapped oncogenes, most cancer-associated rearrangements, evolutionary rearrangements, meiotic chiasmata, and X-ray-induced breaks. Flavors with extremely Alu-rich flavors are also involved in melphalan-induced rearrangements, pachytene stretching, and mitotic chiasmata. Frequencies of CpG islands, CCGCCC boxes, retroposon families, and genes are characteristic to each chromatin flavor and will facilitate alignment of genome sequences onto ideograms of chromatin flavor. The influence of chromatin flavor on the evolution of a gene's sequence is so strong that one can infer the flavor of the band in which a gene resides from the sequence of the gene itself. Correlation coefficients for many pairs of mapped genetic variables, while globally high, are quite low within bands of one flavor, implicating a concerted mode of evolution for bands of one chromatin flavor.

Biological Evolution↗

Complete primary structure for the zymogen of human complement factor B.

The entire amino acid sequence of complement factor B has been established combining both protein and DNA sequencing strategies. The zymogen consists of 739 amino acids, has four asparagine-linked carbohydrate sites, and has independently disulfide-bonded NH2- and COOH-terminal regions. The catalytic subunit, Bb, is a unique serine protease containing 259 amino acids that are not integral to any of the classical serine proteases. It is proposed that this region of the Bb fragment functions as a cofactor-binding domain for C3b. The Ba fragment was found to contain three regions of internal sequence homology which were unrelated to the "kringle" regions of prothrombin and plasminogen and which suggest an independent evolution for the B genome. Sequence alignment of the active site of B to the serine proteases was made using the three-dimensional structures of chymotrypsin and trypsin as molecular models. Three stretches within the hypothetical model for B contrast markedly with all known serine proteases in both amino acid sequences and predicted configuration. It is suggested that these "altered" regions contribute at least in part to the formation of the catalytic region of the C3 convertase.

Amino Acid Sequence↗

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗

Computational prediction of cis-regulatory modules from multispecies alignments using Galaxy, Table Browser, and GALA.

One major goal of genomics is to identify all the functional sequences in genomes, including sequences that regulate the expression of genes. Sequence conservation is a good, albeit imperfect, guide to these functional elements. We describe how to use publicly available servers (Galaxy, the UCSC Table Browser, and GALA) to find genomic sequences whose alignments (from blastZ and multiZ) show properties associated with cis-regulatory modules, such as high conservation score, high regulatory potential score, and conserved transcription factor binding sites. Links to these servers can be accessed at http://www.bx.psu.edu/ and http://genome.ucsc.edu/.

Animals↗