Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Selecting for functional alternative splices in ESTs.

The expressed sequence tag (EST) collection in dbEST provides an extensive resource for detecting alternative splicing on a genomic scale. Using genomically aligned ESTs, a computational tool (TAP) was used to identify alternative splice patterns for 6400 known human genes from the RefSeq database. With sufficient EST coverage, one or more alternatively spliced forms could be detected for nearly all genes examined. To identify high (>95%) confidence observations of alternative splicing, splice variants were clustered on the basis of having mutually exclusive structures, and sample statistics were then applied. Through this selection, alternative splices expected at a frequency of >5% within their respective clusters were seen for only 17%-28% of genes. Although intron retention events (potentially unspliced messages) had been seen for 36% of the genes overall, the same statistical selection yielded reliable cases of intron retention for <5% of genes. For high-confidence alternative splices in the human ESTs, we also noted significantly higher rates both of cross-species conservation in mouse ESTs and of validation in the GenBank mRNA collection. We suggest quantitative analytical approaches such as these can aid in selecting useful targets for further experimental characterization and in so doing may help elucidate the mechanisms and biological implications of alternative splicing.

Alternative Splicing↗

Similarity and differences in the Lactobacillus acidophilus group identified by polyphasic analysis and comparative genomics.

A set of lactobacilli were investigated by polyphasic analysis. Multilocus sequence analysis, DNA typing, microarray analysis, and in silico whole-genome alignments provided a remarkably consistent pattern of similarity within the Lactobacillus acidophilus complex. On microarray analysis, 17 and 5% of the genes from Lactobacillus johnsonii strain NCC533 represented variable and strain-specific genes, respectively, when tested against four independent isolates of L. johnsonii. When projected on the NCC533 genome map, about 10 large clusters of variable genes were identified, and they were enriched around the terminus of replication. A quarter of the variable genes and two-thirds of the strain-specific genes were associated with mobile DNA. Signatures for horizontal gene transfer and modular evolution were found in prophages and in DNA from the exopolysaccharide biosynthesis cluster. On microarray hybridizations, Lactobacillus gasseri strains showed a shift to significantly lower fluorescence intensities than the L. johnsonii test strains, and only genes encoding very conserved cellular functions from L. acidophilus hybridized to the L. johnsonii array. In-silico comparative genomics showed extensive protein sequence similarity and genome synteny of L. johnsonii with L. gasseri, L. acidophilus, and Lactobacillus delbrueckii; moderate synteny with Lactobacillus casei; and scattered X-type sharing of protein sequence identity with the other sequenced lactobacilli. The observation of a stepwise decrease in similarity between the members of the L. acidophilus group suggests a strong element of vertical evolution in a natural phylogenetic group. Modern whole-genome-based techniques are thus a useful adjunct to the clarification of taxonomical relationships in problematic bacterial groups.

Genetic Variation↗

Theatre: A software tool for detailed comparative analysis and visualization of genomic sequence.

Theatre is a web-based computing system designed for the comparative analysis of genomic sequences, especially with respect to motifs likely to be involved in the regulation of gene expression. Theatre is an interface to commonly used sequence analysis tools and biological sequence databases to determine or predict the positions of coding regions, repetitive sequences and transcription factor binding sites in families of DNA sequences. The information is displayed in a manner that can be easily understood and can reveal patterns that might not otherwise have been noticed. In addition to web-based output, Theatre can produce publication quality colour hardcopies showing predicted features in aligned genomic sequences. A case study using the p53 promoter region of four mammalian species and two fish species is described. Unlike the mammalian sequences the promoter regions in fish have not been previously predicted or characterized and we report the differences in the p53 promoter region of four mammals and that predicted for two fish species. Theatre can be accessed at http://www.hgmp.mrc.ac.uk/Registered/Webapp/theatre/.

Animals↗

Variation resources at UC Santa Cruz.

The variation resources within the University of California Santa Cruz Genome Browser include polymorphism data drawn from public collections and analyses of these data, along with their display in the context of other genomic annotations. Primary data from dbSNP is included for many organisms, with added information including genomic alleles and orthologous alleles for closely related organisms. Display filtering and coloring is available by variant type, functional class or other annotations. Annotation of potential errors is highlighted and a genomic alignment of the variant's flanking sequence is displayed. HapMap allele frequencies and linkage disequilibrium (LD) are available for each HapMap population, along with non-human primate alleles. The browsing and analysis tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Alleles↗

snoSeeker: an advanced computational package for screening of guide and orphan snoRNA genes in the human genome.

Small nucleolar RNAs (snoRNAs) represent an abundant group of non-coding RNAs in eukaryotes. They can be divided into guide and orphan snoRNAs according to the presence or absence of antisense sequence to rRNAs or snRNAs. Current snoRNA-searching programs, which are essentially based on sequence complementarity to rRNAs or snRNAs, exist only for the screening of guide snoRNAs. In this study, we have developed an advanced computational package, snoSeeker, which includes CDseeker and ACAseeker programs, for the highly efficient and specific screening of both guide and orphan snoRNA genes in mammalian genomes. By using these programs, we have systematically scanned four human-mammal whole-genome alignment (WGA) sequences and identified 54 novel candidates including 26 orphan candidates as well as 266 known snoRNA genes. Eighteen novel snoRNAs were further experimentally confirmed with four snoRNAs exhibiting a tissue-specific or restricted expression pattern. The results of this study provide the most comprehensive listing of two families of snoRNA genes in the human genome till date.

Algorithms↗

Neighboring base composition and transversion/transition bias in a comparison of rice and maize chloroplast noncoding regions.

The correspondence between the transversion/transition ratio and the neighboring base composition in chloroplast DNA is examined. For 18 noncoding regions of the chloroplast genome, alignments between rice (Oryza sativa) and maize (Zea mays) were generated by two different methods. Difficulties of aligning noncoding DNA are discussed, and the alignments are analyzed in a manner that reduces alignment artifacts. Sequence divergence is < 10%, so multiple substitutions at a site are assumed to be rare. Observed substitutions were analyzed with respect to the A+T content of the two immediately flanking bases. It is shown that as this content increases, the proportion of transversions also increases. When both the 5'- and 3'-flanking nucleotides are G or C (A+T content of 0), only 25% of the observed substitutions are transversions. However, when both the 5'- and 3'-flanking nucleotides are A or T (A+T content of 2), 57% of the observed substitutions are transversions. Therefore, the influence of flanking base composition on substitutions, previously reported for a single noncoding region, is a general feature of the chloroplast genome.

Algorithms↗

Intact {alpha}-1,2-endomannosidase is a typical type II membrane protein.

Rat endomannosidase is a glycosidic enzyme that catalyzes the cleavage of di-, tri-, or tetrasaccharides (Glc(1-3)Man), from N-glycosylation intermediates with terminal glucose residues. To date it is the only characterized member of this class of endomannosidic enzymes. Although this protein has been demonstrated to localize to the Golgi lumenal membrane, the mechanism by which this occurs has not yet been determined. Using the rat endomannosidase sequence, we identified three homologs, one each in the human, mouse, and rat genomes. Alignment of the four encoded protein sequences demonstrated that the newly identified sequences are highly conserved but differed significantly at the N-terminus from the previously reported protein. In this study we have cloned two novel endomannosidase sequences from rat and human cDNA libraries, but were unable to amplify the open reading frame of the previously reported rat sequence. Analysis of the rat genome confirmed that the 59- and 39-termini of the previously reported sequence were in fact located on different chromosomes. This, in combination with our inability to amplify the previously reported sequence, indicated that the N-terminus of the rat endomannosidase sequence previously published was likely in error (a cloning artifact), and that the sequences reported in the current study encode the intact proteins. Furthermore, unlike the previous sequence, the three ORFs identified in this study encode proteins containing a single N-terminal transmembrane domain. Here we demonstrate that this region is responsible for Golgi localization and in doing so confirm that endomannosidase is a type II membrane protein, like the majority of other secretory pathway glycosylation enzymes.

Amino Acid Sequence↗

Compositional evolution of noncoding DNA in the human and chimpanzee genomes.

We have examined the compositional evolution of noncoding DNA in the primate genome by comparison of lineage-specific substitutions observed in 1.8 Mb of genomic alignments of human, chimpanzee, and baboon with 6542 human single-nucleotide polymorphisms (SNPs) rooted using chimpanzee sequence. The pattern of compositional evolution, measured in terms of the numbers of GC-->AT and AT-->GC changes, differs significantly between fixed and polymorphic sites, and indicates that there is a bias toward fixation of AT-->GC mutations, which could result from weak directional selection or biased gene conversion in favor of high GC content. Comparison of the frequency distributions of a subset of the SNPs revealed no significant difference between GC-->AT and AT-->GC polymorphisms, although AT-->GC polymorphisms in regions of high GC segregate at slightly higher frequencies on average than GC-->AT polymorphisms, which is consistent with a fixation bias favoring high GC in these regions. However, the substitution data suggest that this fixation bias is relatively weak, because the compositional structure of the human and chimpanzee genomes is becoming homogenized, with regions of high GC decreasing in GC content and regions of low GC increasing in GC content. The rate and pattern of nucleotide substitution in 333 Alu repeats within the human-chimpanzee-baboon alignments are not significantly affected by the GC content of the region in which they are inserted, providing further evidence that, since the time of the human-chimpanzee ancestor, there has been little or no regional variation in mutation bias.

Alleles↗

Reconstruction of ancestral nucleotide sequences and estimation of substitution frequencies in a star phylogeny.

Maximum likelihood phylogeny reconstruction methods are widely used in uncovering and assessing the evolutionary history and relationships of natural systems. However, several simplifying assumptions commonly made in this analysis limit the explanatory power of the results obtained. We present an algorithm that performs the phylogenetic analysis without making the common assumptions for sequence data from at least three leaf nodes in a star phylogeny. In particular, the underlying nucleotide substitution model does not have to be reversible and may include neighbor-dependent processes like the CpG methylation deamination process (CpG-effect). The base composition of the sequences at the external nodes and the one of the ancestral sequence may be different from each other and they do not have to be stationary state distributions of the corresponding substitution model. The algorithm is able to reconstruct the ancestral base composition and accurately estimate substitution frequencies in the branches of the star phylogeny. Extensive tests on simulated data validate the very favorable performance of the algorithm. As an application we present the analysis of aligned genomic sequences from human, mouse, and dog. Different substitution pattern can be observed in the three lineages.

Algorithms↗

Strand bias in complementary single-nucleotide polymorphisms of transcribed human sequences: evidence for functional effects of synonymous polymorphisms.

BACKGROUND: Complementary single-nucleotide polymorphisms (SNPs) may not be distributed equally between two DNA strands if the strands are functionally distinct, such as in transcribed genes. In introns, an excess of A<-->G over the complementary C<-->T substitutions had previously been found and attributed to transcription-coupled repair (TCR), demonstrating the valuable functional clues that can be obtained by studying such asymmetry. Here we studied asymmetry of human synonymous SNPs (sSNPs) in the fourfold degenerate (FFD) sites as compared to intronic SNPs (iSNPs). RESULTS: The identities of the ancestral bases and the direction of mutations were inferred from human-chimpanzee genomic alignment. After correction for background nucleotide composition, excess of A-->G over the complementary T-->C polymorphisms, which was observed previously and can be explained by TCR, was confirmed in FFD SNPs and iSNPs. However, when SNPs were separately examined according to whether they mapped to a CpG dinucleotide or not, an excess of C-->T over G-->A polymorphisms was found in non-CpG site FFD SNPs but was absent from iSNPs and CpG site FFD SNPs. CONCLUSION: The genome-wide discrepancy of human FFD SNPs provides novel evidence for widespread selective pressure due to functional effects of sSNPs. The similar asymmetry pattern of FFD SNPs and iSNPs that map to a CpG can be explained by transcription-coupled mechanisms, including TCR and transcription-coupled mutation. Because of the hypermutability of CpG sites, more CpG site FFD SNPs are relatively younger and have confronted less selection effect than non-CpG FFD SNPs, which can explain the asymmetric discrepancy of CpG site FFD SNPs vs. non-CpG site FFD SNPs.

Algorithms↗

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160&#x2009;myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals↗

Spontaneous deletions and duplications of sequences in the genome of cowpox virus.

Examination of the genomes of 10 white-pock variants of cowpox virus strain Brighton red (CPV-BR) revealed that 9 of them had lost 32 to 38 kilobase pairs (kbp) from their right-hand ends and that the deleted sequences had been replaced by inverted copies of regions from 21 to 50 kbp long from the left-hand end of the genome. These variants thus possess inverted terminal repeats (ITRs) from 21 to 50 kbp long; all are longer than the ITRs of CPV-BR (10 kbp). The 10th variant is a simple deletion mutant that has lost the sequences between 32 and 12 kbp from the right-hand end of the genome. The limits of the inner ends of the observed deletions (between 32 and 38 kbp from the right-hand end of the CPV-BR genome) appear to be defined by the location of the nearest essential gene on the one hand and the location of the gene that encodes "pock redness" on the other. The genomes of the deletion/duplication white-pock variants appear to have been generated either by single crossover recombinational events between two CPV-BR genomes aligned in opposite directions or by the nonreciprocal transfer of genetic information. The sites where such recombination/transfer occurred were sequenced in four variants. In all of them, the sequences adjacent to such sites show no sequence homology or any other unusual structural feature. The analogous sites at the internal ends of the two ITRs of CPV-BR also were sequenced and also show no unusual features. It is likely that the ITRs of CPV-BR and of its white-pock variants, and probably those of other orthopox-virus genomes, arise as a result of nonhomologous recombination or by random nonreciprocal transfer of genetic information.

Base Sequence↗

An isoleucyl-tRNA synthetase gene from Campylobacter jejuni.

A complete isoleucyl-tRNA synthetase gene (ileS) of Campylobacter jejuni was isolated from a C. jejuni TGH9011 genomic DNA library constructed in pBluescript. The complete coding sequence, flanking regions and transcription start point were determined. The deduced isoleucyl-tRNA synthetase (IleRS) had 917 amino acids with a molecular mass of 105,399 Da, which was consistent with the observed size of 105 kDa in Escherichia coli maxicells. The ileS gene was mapped onto the physical map of the C. jejuni genome. Alignment of the C. jejuni IleRS sequence with six other bacterial IleRS sequences and two lower eukaryotic IleRS sequences identified seven conserved motifs, including the two signature sequences, HIGH and KMSKS, of class I aminoacyl-tRNA synthetases.

Amino Acid Sequence↗

(Re)imagining the Future of Genetic Counseling: A Reflexive Qualitative Analysis of Sociopolitical Power, Cultural Safety, Systemic Racism, and Comparative Practice in the United Kingdom, Aotearoa New Zealand and, Australia.

Genetic counseling is undergoing a rapid transformation as genomic medicine becomes embedded within mainstream healthcare systems. At the same time, the profession is being challenged to respond to systemic racism, colonial legacies, technological change, and evolving expectations regarding equity and justice. Historically, genetic counseling emerged within twentieth-century medical genetics and was influenced by political, social, scientific, and medical forces that included eugenic ideology, values, and practices. The profession has since evolved substantially toward psychosocial, patient-centered, and non-directive models of care. Contemporary debates regarding "newgenics" or "neugenics" further demonstrate how concerns regarding equity, reproductive ethics, disability, and genomic stratification continue to shape genomic healthcare discourse. This qualitative reflexive practice paper explores how systemic racism, colonial legacy, cultural safety and structural power shape genetic counseling practice in the United Kingdom (UK), Aotearoa New Zealand and Australia, and how these forces continue to reshape the profession's future identity. A reflexive, narrative, and comparative qualitative approach was employed, grounded in the authors' lived professional experiences across UK and Australasian contexts and informed by purposively selected policy, professional and scholarly literature relating to cultural safety, dignity, anti-racism, and Human Rights-Based Decision-Making. Through iterative reflexive dialogue, comparative analysis, and thematic synthesis, four interrelated themes were developed examining sociopolitical context, systemic racism, cultural safety and technologization within contemporary genetic counseling practice. Comparative analysis identified substantial differences in how culturally responsive practice is conceptualized and operationalized across settings. In Aotearoa, cultural safety is strongly shaped by Te Tiriti o Waitangi, bicultural accountability, and M&#x101;ori sovereignty frameworks. In Australia, culturally safer genomic care has increasingly developed through Indigenous-led initiatives and workforce reform, including the Australian Alliance for Indigenous Genomics (ALIGN). In contrast, UK practice remains largely situated within equality, diversity, and inclusion (EDI) frameworks that may insufficiently address systemic racism and structural power within increasingly diverse populations. Reflexive clinical examples demonstrated how inequities may emerge through undocumented patient values, standardized pathways, assumptions regarding autonomy, and misinterpretation of culturally specific communication styles. Re-imagining the future of genetic counseling requires more than just technological advancement. It requires reflexive engagement with dignity, inequity, and the sociopolitical realities of the populations served. These insights re-imagine a culturally grounded, socially responsive future for genetic counseling in an era shaped by genomic mainstreaming, digital transformation, artificial intelligence and workforce reform and one in which the profession remains ethically anchored, relationally attuned, and committed to justice-oriented practice.

Humans↗

PACdb: PolyA Cleavage Site and 3'-UTR Database.

UNLABELLED: The PolyA Cleavage Site and 3'-UTR Database (PACdb) is a web-accessible database that catalogs putative 3'-processing sites and 3'-UTR sequences for multiple organisms. Sites have been identified primarily via expressed sequence tag-genome alignments, enabling delineation of both the specificities and heterogeneity of 3'-processing events. AVAILABILITY: By web browser or CGI: PACdb: http://harlequin.jax.org/pacdb/; AtPACdb: http://harlequin.jax.org/atpacdb/. SUPPLEMENTARY INFORMATION: Available online at http://harlequin.jax.org/pacdb/supplemental.php.

3' Untranslated Regions↗

CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.

SUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards.

Cholera↗

The lacrimal gland transcriptome is an unusually rich source of rare and poorly characterized gene transcripts.

PURPOSE: To sequence and comprehensively analyze human and mouse lacrimal gland transcriptomes as part of the NEIBank project. METHODS: cDNA libraries generated from normal human and mouse lacrimal glands were sequenced and analyzed by PHRED, RepeatMasker, BLAST, and GRIST. Human "lacrimal-preferred genes" and putative gene regulatory elements were respectively identified in UniGene and ConSite, and gene clustering was analyzed by chromosomal mapping. "Hypothetical proteins," identified by keyword search, were verified by genomic alignment and queried in the Conserved Domain database and GEO Profiles. RESULTS: The top six transcripts in human and mouse differed, revealing a previously unappreciated molecular divergence. The human transcriptome is enriched with transcripts from 29 lacrimal-preferred genes and a content of poorly characterized hypothetical proteins, proportionally greater than in all other tissues. Only 45% of lacrimal preferred, but 71% of hypotheticals, have mouse orthologs. Many of the latter display apparently altered cancer expression in the CGAP SAGE library collection-often in keeping with predicted WD40, protein kinase, Src homology 2 and 3, RhoGEF, and pleckstrin homology domains involved in cell signaling. At the genomic level, lacrimal-expressed genes show some evidence of clustering, particularly on human chromosomes 9 and 12. Binding sites for TFAP2A, FOXC1, and other transcription factors are predicted. CONCLUSIONS: Interspecies divergence cautions against use of mouse models of human dry eye syndromes. Lacrimal preferred and hypothetical proteins, gene clustering, and putative gene regulatory elements together provide new clues for a molecular understanding of lacrimal gland function and mechanisms of coordinated tissue-specific transcriptional regulation.

Aged↗

Extensive amplification and transposition of a novel repetitive element, xstir, together with its terminal inverted repeat in the evolution of Xenopus.

A DNA fragment containing short tandem repeat sequences (approximately 86-bp repeat) was isolated from a Xenopus laevis cDNA library. Southern blot and in situ hybridization analyses revealed that the repeat was highly dispersed in the genome and was present at approximately 1 million copies per haploid genome. We named this element Xstir (Xenopus short tandemly and invertedly repeating element) after its arrangement in the genome. The majority of the genomic Xstir sequences were digested to monomer and dimer sizes with several restriction enzymes. Their sequences were found to be highly homogeneous and organized into tandem arrays in the genome. Alignment analyses of several known sequences showed that some of the Xstir-like sequences were also organized into interspersed inverted repeats. The inverted repeats consisted of an inverted pair of two differently modified Xstirs separated by a short insert. In addition, these were framed by another novel inverted repeat (Xstir-TIR). The Xstir-TIR sequence was also found at the ends of tandem Xstir arrays. Furthermore, we found that Xstir-TIR was linked to a motif characterizing the T2 family which belonged to a vertebrate MITE (miniature inverted-repeat transposable element) family, suggesting the importance of Xstir-TIR for their amplification and transposition. The present study of 11 anuran and 2 urodele species revealed that Xstir or Xstir-like sequences were extensively amplified in the three Xenopus species. Genomic Xstir populations of X. borealis and X. laevis were mutually indistinguishable but significantly different from that of X. tropicalis.

Animals↗