Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Phylogeny, sequence conservation, and functional complementation of the SBDS protein family.

The Shwachman-Bodian-Diamond syndrome (SBDS) protein family occurs widely in nature, although its function has not been determined. Comprehensive database searches revealed SBDS homologues from 159 species, including examples from all sequenced archaeal and eukaryotic genomes and all eukaryotic kingdoms. Sequence alignment with ClustalX and MUSCLE algorithms led to the identification of conserved residues that occurred predominantly in the amino-terminal FYSH domain where they appeared to contribute to protein folding or stability. Only SBDS residue Gly91 was invariant in all species. Four distantly related protists were found to have two divergent SBDS genes in their genomes. In each case, phylogenetic analyses and the identification of shared sequence features suggested that one gene was derived from lateral gene transfer. We also identified a shared C-terminal zinc finger domain fusion in flowering plants and chromalveolates that may shed light on the function of the protein family and the evolutionary histories of these kingdoms. To assess the extent of SBDS functional conservation, we carried out complementation studies of SBDS homologues and interspecies chimeras in Saccharomyces cerevisiae. We determined that the FYSH domain was widely interchangeable among eukaryotes, while domain 2 imparted species specificity to protein function. Domain 3 was largely dispensable for function in our yeast complementation assay. Overall, the phylogeny of SBDS was shared with a group of proteins that were markedly enriched for RNA metabolism and/or ribosome-associated functions. These findings link Shwachman-Diamond syndrome to other bone marrow failure syndromes with defects in nucleolus-associated processes, including Diamond-Blackfan anemia, cartilage-hair hypoplasia, and dyskeratosis congenita.

Amino Acid Sequence↗

Pancreatic triacylglycerol lipase in a hibernating mammal. I. Novel genomic organization.

Pancreatic triacylglycerol lipase (PTL) is expressed in novel locations during hibernation in the thirteen-lined ground squirrel (Spermophilus tridecemlineatus). PTL cDNAs isolated from two of these locations, heart and white adipose tissue (WAT), contain divergent 5'-untranslated regions (5'-UTRs) suggesting alternative promoter usage or the possibility of multiple PTL genes in the ground squirrel genome. In addition, cDNAs isolated from WAT contain tracts of retroviral sequence in their 5'-UTRs. Our examination of PTL genomic clones isolated from a thirteen-lined ground squirrel genomic DNA library, coupled with genomic Southern blot analysis, enabled us to conclude that PTL mRNAs expressed in heart and WAT are the products of the same single-copy gene. The 5' portion of this gene spans 9.2 kb, is composed of 6 exons, and contains a full-length endogenous retroviral genome with conserved long terminal repeats (LTRs). Alignment of the ground squirrel PTL gene with the mouse, rat, and human PTL genes indicates that this retrovirus inserted into the ground squirrel genome approximately 200 bases upstream of the original PTL transcriptional start site. The insertion is a relatively recent event based on largely intact open-reading frames containing minimal frame-shift and nonsense mutations. The high-percentage identity (99.2%) shared between the 5'- and 3'-LTRs of this endogenous retrovirus suggests that the insertion occurred as recently as 300,000 years ago.

Adipose Tissue↗

Comparative genomics and functional roles of the ATP-dependent proteases Lon and Clp during cytosolic protein degradation.

The general pathway involving adenosine triphosphate (ATP)-dependent proteases and ATP-independent peptidases during cytosolic protein degradation is conserved, with differences in the enzymes utilized, in organisms from different kingdoms. Lon and caseinolytic protease (Clp) are key enzymes responsible for the ATP-dependent degradation of cytosolic proteins in Escherichia coli. Orthologs of E. coli Lon and Clp were searched for, followed by multiple sequence alignment of active site residues, in genomes from seventeen organisms, including representatives from eubacteria, archaea, and eukaryotes. Lon orthologs, unlike ClpP and ClpQ, are present in most organisms studied. The roles of these proteases as essential enzymes and in the virulence of some organisms are discussed.

Adenosine Triphosphate↗

Genomic organization of a potential human DNA-crosslink repair gene, KIAA0086.

In the present study, we describe the genomic structure of the KIAA0086 gene and the 5'-flanking sequence. The analysis is based on the alignment of the KIAA0086 cDNA and a corresponding genomic BAC sequence which was identified in a basic BLAST similarity search using the cDNA sequence as a template. The gene contains nine exons spanning approximately 20 kb. All splice sites conform to the GT-AG rule. Analysis of the upstream untranscribed region identified one GC box but no TATA box, suggesting that the KIAA0086 gene is a housekeeping gene. The promoter region contains putative recognition sites for several transcription factors, e.g., AP1, Sp1 and NFkappaB. The homology of the KIAA0086 gene to the yeast SNM1 gene, which is involved in the cellular response to DNA-interstrand crosslinks, is discussed with respect to a possible role of the KIAA0086 gene in the human disorder, Fanconi anemia.

Amino Acid Sequence↗

Detection of human sapovirus by real-time reverse transcription-polymerase chain reaction.

Sapovirus (SaV) is an agent of gastroenteritis for humans and swine, and is divided into five distinct genogroups (GI-GV) based on its capsid gene sequences. Typical methods of SaV detection include electron microscopy (EM), enzyme-linked immunosorbent assay (ELISA), and reverse transcription-polymerase chain reaction (RT-PCR). A novel TaqMan-based real-time RT-PCR assay was developed that is sensitive and has the ability to detect the broad range of genetically diverse human SaV strains. A nucleotide alignment of 10 full-length SaV genome sequences was subjected to similarity plot analysis, which indicated that the most conserved site was the polymerase-capsid junction in open reading frame 1 (ORF1). Based on multiple alignments of the 27 available sequences encoding this junction, we designed sets of primers and TaqMan MGB probes that detect human SaV GI, GII, GIV, and GV sequences in a single tube. The reactivity was confirmed with SaV GI, GII, GIV, and GV control plasmids, and the efficiency ranged from 2.5 x 10(7) to 2.5 x 10(1) copies per tube. Analysis using clinical stool specimens revealed that the present system was capable of detecting SaV GI, GII, GIV, and GV sequences, and no cross-reactivity was observed against other enteric viruses, including norovirus (NoV), rotavirus, astrovirus, and adenovirus. This is the first real-time RT-PCR system that could detect all genogroups of human sapoviruses.

Australia↗

FISH digital imaging microscopy in mosquito genomics.

The yellow fever mosquito, Aedes aegypti, transmits pathogens that affect both humans and livestock, and has been the focus of extensive research to identify genetic loci that may be useful in control strategies. Fluorescence in situ hybridization (FISH) and digital imaging microscopy have provided a rapid mechanism to populate the physical map with probes derived from genetic markers, cDNAs and recombinant genomic libraries. When the physical and genetic linkage maps are aligned, map-based cloning will allow the rapid isolation of target genomic sequences. The strategy of FISH mapping and the results of initial hybridization studies are reviewed here by Martin Ferguson, Susan Brown and Dennis Knudson. An Ae. aegypti-specific genomic database, which collates data from mapping studies, sequences, references and other relevant information, is also discussed.

Journal Article↗

Molecular characterization of an isolate of citrus tristeza virus that causes severe symptoms in sweet orange.

The complete sequence (19,249 nucleotides) of the genome of citrus tristeza virus (CTV) isolate SY568 was determined. The genome organization is identical to that of the previously determined CTV-T36 and CTV-VT isolates. Sequence comparisons revealed that CTV-SY568, a severe stem-pitting isolate from California, has more than 87% overall sequence identity with CTV-VT, a seedling yellows isolate from Israel. Although SY568 has an overall sequence identity of 81% with CTV-T36, a quick decline isolate from Florida, the sequence identity in the 3' half of the genome is over 90% while the sequence identity in the 5' half of the genome is as low as 56%. Based on the sequence alignments of these three isolates, sequences in the 3' half of the genome are generally well conserved, while the sequences in the 5' half are relatively divergent. Sequence data of independent overlapping clones from the CTV-SY568 genome revealed two regions with highly divergent sequences. In open reading frame 1b (RNA dependent RNA polymerase), there were 118 nucleotide differences that lead to 16 amino acid changes. In the open reading frame of the divergent coat protein gene, 5 amino acid changes result from 48 nucleotide differences. Most differences occurred in the third position of the codons, and resulted in silent amino acid substitutions. RNase protection assays demonstrated that most of the clones obtained are representative of the major RNA species of this isolate. Northern analysis indicated that CTV-SY568 accumulated more viral RNA including genomic and certain subgenomic RNAs than isolates VT or T36 in sweet orange.

Base Sequence↗

Whole-genome characterization and phylogenetic placement of Fusarium oxysporum f. sp. vasinfectum isolates.

Fusarium wilt of cotton, caused by Fusarium oxysporum f. sp. vasinfectum (Fov), remains a persistent threat to cotton production worldwide. Among the known races, Fov race 4 and its extra-virulent variants cause particularly severe losses in Upland cotton. Although several Fov genome assemblies have been assigned to races, the genomic diversity and evolutionary relationships among pathogenic and non-pathogenic isolates associated with cotton outbreaks remain poorly understood at the whole-genome level. This study addressed these gaps by generating and comparing high-quality genome assemblies of four Fusarium isolates collected from Texas cotton fields: two pathogenic (TX17-24 and TX18-9) and two non-pathogenic (TX17-6 and TX18-6). Draft assemblies were generated using Oxford Nanopore long reads and polished with Illumina reads. Comparative genomic analyses showed that pathogenic isolates possessed larger genomes and more conserved orthologous families, whereas non-pathogenic isolates contained more unique genes. Analyses of predicted secreted effectors, transposable elements, and carbohydrate-active enzymes further distinguished pathogenic and non-pathogenic lineages, suggesting roles in virulence adaptation and genome plasticity. Phylogenomic analyses using k-mer-based, assembly- and alignment-free methods incorporated all available long-read Fov genomes and revealed substantial genetic diversity within races 1 and 4, clustering isolates into multiple sublineages. These findings show that Fov race diversification is underestimated when based on traditional classification schemes and may be shaped by host specialization, geographic separation, or horizontal gene transfer. This work advances our understanding of the genomic diversity and evolutionary dynamics of Fov and establishes a foundation for improved race identification and characterization of Fusarium wilt pathogenesis in cotton.

Fusarium oxysporum↗

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article↗

Identification and analysis of the human neural polypyrimidine tract binding protein (nPTB) gene promoter region.

Neural polypyrimidine tract-binding protein nPTB, originally identified as the neuronal counterpart of the hnRNPI/PTB protein, is an RNA binding protein involved into tissue-specific alternative splicing regulation. Here we describe the functional characterization of the promoter sequence of nPTB in HeLa and neuroblastoma cells. By means of genomic sequence analysis we have isolated and cloned a 1587-base pair region upstream the human nPTB coding region. The nPTB proximal promoter, although rich in G+C content and presenting putative binding sites for the transcription factors Sp1, NF-1, NF-kB and Oct-1, lacks a typical TATA box. Luciferase transient expression assays using deletion mutants have identified the proximal promoter region at -125 relative to the transcription start site. Alignment of human, murine and chimpanzee genomic sequences upstream the nPTB exon 1 has provided evidences for the evolutionary conservation of specific transcription factor binding sites.

5' Flanking Region↗

Correlation Between Infectivity and qRT-PCR Values for Murine Norovirus Recovered from Frozen Berries.

Human norovirus (HuNoV) is the leading cause of acute gastroenteritis globally, with frozen berries frequently implicated in foodborne outbreaks. Current surveillance relies on quantitative reverse transcription PCR (qRT-PCR), which cannot differentiate between infectious and non-infectious viral particles, complicating risk assessment. This study is aimed to establish the minimum viral load on frozen berries detectable by qRT-PCR that corresponds to infectious virus, using murine norovirus (MNV) as a surrogate for HuNoV. Frozen raspberries were artificially inoculated with serial dilutions of MNV (7.1-1.0 log PFU/25 g) and processed using the ISO 15216:2017 method. Infectious virus was quantified by plaque assay, and viral RNA was detected by qRT-PCR. The limit of detection (LOD) for cell culture was 3.1 log PFU/25 g, whereas qRT-PCR extended sensitivity to 1.0 log PFU/25 g (Ct value at 36.7 ± 0.6), representing a 2-log difference. Recovery rates for infectious virus exceeded the ISO 15,216 minimum threshold (1%), and PCR inhibition was negligible. We next examined the extraction efficiency for both infectious MNV and its genetic material from frozen strawberries at inoculation levels higher than the LOD, and observed that the viral recovery from frozen strawberries is very similar to viral recovery from frozen raspberries with no significant differences between them. The disparity between LODs indicates that a substantial proportion of MNV genomes detected by qRT-PCR do not represent infectious particles, aligning with previous findings that one PFU may correspond to multiple genome copies. Given that many surveillance studies report high Ct values (> 35), our data suggest that such detections may not indicate viable virus, underscoring the importance of contextualizing qRT-PCR results with epidemiological evidence. These findings highlight the need for cautious interpretation of surveillance data, particularly for public health decision-making.

Norovirus↗

Patterns in spontaneous mutation revealed by human-baboon sequence comparison.

We have analyzed the alignment of a long homologous region of the human and baboon genomes (approximately 1.5 Mb). We show that the frequency of gaps between aligned segments decreases slowly with gap length, indicating that several successive nucleotides are often deleted or inserted in one event. By contrast, runs of consecutive mismatches decrease rapidly in frequency with increasing length, following an exponential distribution, indicating that nucleotides are mostly substituted one at a time. Nucleotide substitutions are clumped at the scales of <10 and 1000-10,000 nucleotides, but show almost no aggregation at the scales of <10-100 and over approximately 50,000 nucleotides. Apparently, two rather different factors make the substitution rate not exactly uniform along the DNA sequence. Comparison of regions of very similar genomes that are approximately selectively neutral makes it possible to study spontaneous mutation at a new level of resolution.

Animals↗

Identification and characterization of the potential promoter regions of 1031 kinds of human genes.

To understand the mechanism of transcriptional regulation, it is essential to identify and characterize the promoter, which is located proximal to the mRNA start site. To identify the promoters from the large volumes of genomic sequences, we used mRNA start sites determined by a large-scale sequencing of the cDNA libraries constructed by the "oligo-capping" method. We aligned the mRNA start sites with the genomic sequences and retrieved adjacent sequences as potential promoter regions (PPRs) for 1031 genes. The PPR sequences were searched to determine the frequencies of major promoter elements. Among 1031 PPRs, 329 (32%) contained TATA boxes, 872 (85%) contained initiators, 999 (97%) contained GC box, and 663 (64%) contained CAAT box. Furthermore, 493 (48%) PPRs were located in CpG islands. This frequency of CpG islands was reduced in TATA(+)/Inr(+) PPRs and in the PPRs of ubiquitously expressed genes. In the PPRs of the CGM2 gene, the DRA gene, and the TM30pl genes, which showed highly colon specific expression patterns, the consensus sequences of E boxes were commonly observed. The PPRs were also useful for exploring promoter SNPs.

Base Sequence↗

Genomics via optical mapping. III: Contiging genomic DNA.

In this paper, we describe our algorithmic approach to constructing an alignment of (contiging) a set of restriction maps created from the images of individual genomic (uncloned) DNA molecules digested by restriction enzymes. Generally, these DNA segments are sized in the range of 1-4 Mb. The goal is to devise contiging algorithms capable of producing high-quality composite maps rapidly and in a scaleable manner. The resulting software is a key component of our physical mapping automation tools and has been used to create complete maps of various microorganisms (E. coli, P. falciparum and D. radiodurans). Experimental results match known sequence data.

Algorithms↗

A genome-wide, end-sequenced 129Sv BAC library resource for targeting vector construction.

The majority of gene-targeting experiments in mice are performed in 129Sv-derived embryonic stem (ES) cell lines, which are generally considered to be more reliable at colonizing the germ line than ES cells derived from other strains. Gene targeting is reliant on homologous recombination of a targeting vector with the host ES cell genome. The efficiency of recombination is affected by many factors, including the isogenicity (H. te Riele et al., 1992, Proc. Natl. Acad. Sci. USA 89, 5128-5132) and the length of homologous sequence of the targeting vector and the location of the target locus. Here we describe the double-end sequencing and mapping of 84,507 bacterial artificial chromosomes (BACs) generated from AB2.2 ES cell DNA (129S7/SvEvBrd-Hprtb-m2). We have aligned these BACs against the mouse genome and displayed them on the Ensembl genome browser, DAS: 129S7/AB2.2. This library has an average insert size of 110.68 kb and average depth of genome coverage of 3.63- and 1.24-fold across the autosomes and sex chromosomes, respectively. Over 97% of the mouse genome and 99.1% of Ensembl genes are covered by clones from this library. This publicly available BAC resource can be used for the rapid construction of targeting vectors via recombineering. Furthermore, we show that targeting vectors containing DNA recombineered from this BAC library can be used to target genes efficiently in several 129-derived ES cell lines.

Animals↗

Sequence diversification of the FK506-binding proteins in several different genomes.

Sequences of FK506-binding proteins (FKBPs) from four genomes of the following organisms were compared: the prokaryote Escherichia coli, the lower eukaryote Saccharomyces cerevisiae, the plant Arabidopsis thaliana, the nematode Caenorhabditis elegans and a composite of 14 unique FKBPs from two mammalian organisms Homo sapiens (man) and Mus musculus (domestic mouse). A singular FK506-like binding domain (FKBD) has about 12 kDa and occurs in the form of archetypal FKBP-12 and as a part of different proteins ranging in size from 13 to 135 kDa. Some organisms may contain a variable number of proteins which consist from two to four consecutively fused FKBDs. In the 12-kDa subgroup of archetypal FKBPs sequence identity (ID) varies from 100 to 83% (mammalian FKBPs-12), 75-50% in mammalian vs. invertebrate FKBPs-12, and fall to about 30% for pairwise sequence comparisons of mammalian and bacterial FKBPs-12 which suggests that their sequences are divergent. Multiple sequence alignment of FKBPs from the four genomes and a set of unique mammalian FKBPs does not contain any explicit consensus sequence but certain sequence positions have conserved physico-chemical characteristics. Variations of hydrophobicity and bulkiness in the multiple sequence alignment are nonsymmetrical because the physico-chemical properties of the aligned sequences changed during evolution. These variations at the sequence positions which are crucial for binding the immunosuppressive macrolide FK506 and peptidyl-prolyl cis/trans isomerase (PPIase) activity are small.

Amino Acid Sequence↗

Antisense transcripts with rice full-length cDNAs.

BACKGROUND: Natural antisense transcripts control gene expression through post-transcriptional gene silencing by annealing to the complementary sequence of the sense transcript. Because many genome and mRNA sequences have become available recently, genome-wide searches for sense-antisense transcripts have been reported, but few plant sense-antisense transcript pairs have been studied. The Rice Full-Length cDNA Sequencing Project has enabled computational searching of a large number of plant sense-antisense transcript pairs. RESULTS: We identified sense-antisense transcript pairs from 32,127 full-length rice cDNA sequences produced by this project and public rice mRNA sequences by aligning the cDNA sequences with rice genome sequences. We discovered 687 bidirectional transcript pairs in rice, including sense-antisense transcript pairs. Both sense and antisense strands of 342 pairs (50%) showed homology to at least one expressed sequence tag other than that of the pair. Microarray analysis showed 82 pairs (32%) out of 258 pairs on the microarray were more highly expressed than the median expression intensity of 21,938 rice transcriptional units. Both sense and antisense strands of 594 pairs (86%) had coding potential. CONCLUSIONS: The large number of plant sense-antisense transcript pairs suggests that gene regulation by antisense transcripts occurs in plants and not only in animals. On the basis of our results, experiments should be carried out to analyze the function of plant antisense transcripts.

DNA, Antisense↗

Hidden diversity in Enterococcus faecalis revealed by CRISPR2 screening: eco-evolutionary insights into a novel subspecies.

Enterococcus faecalis is a commensal bacterium that colonizes the gut of humans and animals and is a major opportunistic pathogen, known for causing multidrug-resistant healthcare-associated infections (HAIs). Its ability to thrive in diverse environments and disseminate antimicrobial resistance genes (ARGs) across ecological niches highlights the importance of understanding its ecological, evolutionary, and epidemiological dynamics. The CRISPR2 locus has been used as a valuable marker for assessing clonality and phylogenetic relationships in E. faecalis. In this study, we identified a group of E. faecalis strains lacking CRISPR2, forming a distinct, well-supported clade. We demonstrate that this clade meets the genomic criteria for classification as a novel subspecies, here referred to as "subspecies B." Through a comprehensive pangenome analysis and comparative genomics, we explored the adaptive ecological traits underlying this diversification process, identifying clade-specific features and their predicted functional roles. Our findings suggest that the frequent isolation of subspecies B from meat products and processing facilities may reflect dissemination routes involving environmental contamination (e.g., water, plants, soil) from avian species. The absence of key virulence traits required for pathogenicity in mammals, particularly humans, and the lack of clinically relevant resistance determinants indicate that subspecies B currently poses minimal threat to public health compared with the broadly disseminated "subspecies A." Nevertheless, the unclear potential for genetic exchange between these subspecies and the frequent association of subspecies B with food sources calls for continued genomic surveillance of E. faecalis from a One Health perspective to detect and mitigate the emergence of high-risk variants in advance.IMPORTANCEExploring intraspecific genetic variability in generalist bacteria with pathogenic potential, such as Enterococcus faecalis, is a key to uncovering stable evolutionary trends. By screening the CRISPR2 locus across a representative set of genomes from diverse sources, this study reveals a previously unrecognized lineage within the population structure of E. faecalis, associated with underexplored nonhuman and nonhospital reservoirs. These findings broaden our knowledge of the species' genetic landscape and shed light on its adaptive strategies and patterns of ecological dissemination. By bridging phylogenetic patterns with variation in genetic defense systems and accessory traits, the study generates testable hypotheses about the genomic determinants and corresponding selective pressures that shape the species' behavior and long-term dissemination. This work offers new perspectives on the eco-evolutionary dynamics of E. faecalis and highlights the value of genomic surveillance beyond clinical settings, in alignment with One Health principles.

Enterococcus faecalis↗