Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference genome”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

An improved probability mapping approach to assess genome mosaicism.

BACKGROUND: Maximum likelihood and posterior probability mapping are useful visualization techniques that are used to ascertain the mosaic nature of prokaryotic genomes. However, posterior probabilities, especially when calculated for four-taxon cases, tend to overestimate the support for tree topologies. Furthermore, because of poor taxon sampling four-taxon analyses suffer from sensitivity to the long branch attraction artifact. Here we extend the probability mapping approach by improving taxon sampling of the analyzed datasets, and by using bootstrap support values, a more conservative tool to assess reliability. RESULTS: Quartets of orthologous proteins were complemented with homologs from selected reference genomes. The mapping of bootstrap support values from these extended datasets gives results similar to the original maximum likelihood and posterior probability mapping. The more conservative nature of the plotted support values allows to focus further analyses on those protein families that strongly disagree with the majority or plurality of genes present in the analyzed genomes. CONCLUSION: Posterior probability is a non-conservative measure for support, and posterior probability mapping only provides a quick estimation of phylogenetic information content of four genomes. This approach can be utilized as a pre-screen to select genes that might have been horizontally transferred. Better taxon sampling combined with subtree analyses prevents the inconsistencies associated with four-taxon analyses, but retains the power of visual representation. Nevertheless, a case-by-case inspection of individual multi-taxon phylogenies remains necessary to differentiate unrecognized paralogy and shared phylogenetic reconstruction artifacts from horizontal gene transfer events.

Bacterial Proteins↗

Strain variability among Kaposi sarcoma-associated herpesvirus (human herpesvirus 8) genomes: evidence that a large cohort of United States AIDS patients may have been infected by a single common isolate.

Previous analysis of the majority of Kaposi's sarcoma (KS) tumors, in both AIDS and non-AIDS populations, has revealed the consistent presence of two small subsegments (open reading frame 25/26 [ORF25/26] and ORF75) of a novel human gamma class herpesvirus genome referred to as KSHV or HHV-8. We have carried out DNA sequence comparisons with DNAs encompassing a total of 2,500 bp each over three separate PCR-amplified fragments from KS lesions and body cavity-based lymphoma (BCBL) samples from 12 distinct patients, including four African and two classical or endemic non-AIDS KS samples. The results revealed differences at 37 of 2,500 nucleotide positions (i.e., 1.5% overall variation). However, the 12 HHV-8 genomes examined fell into three distinct but very narrow subgroupings (A, B, and C strains). All A strain isolates differed from B strain isolates at 16 positions, but of the eight U.S. samples tested, six were A strains, and these differed at no more than two positions among them. Similarly, three of the four African samples were B strains, which differed from each other at only one position. The two C strain genomes also displayed only one nucleotide variation, but they differed from all A strains at 26 positions and from all B strains at 20 positions. One C strain genome was present in all six independent lesions from an AIDS KS patient with disseminated disease, and the other represented a mosaic A/C recombinant genome from the HBL6 cell line derived from a BCBL tumor. Evaluation of previous data suggests that B and C strains may predominate in Africa and that A strains predominate in classical Mediterranean samples. Although both B and C strains are represented in U.S. AIDS patients, the majority (70 to 80%) of samples from the mid-East Coast region at least appear to be virtually identical, supporting the concept that they may all derive from the spread during the AIDS epidemic of a single recently transmitted infectious agent.

AIDS-Related Opportunistic Infections↗

Transposable elements create distinct genomic niches for effector evolution among Magnaporthe oryzae lineages.

BACKGROUND: Plant-pathogen interactions are characterized by evolutionary arms races. At the molecular level, fungal effectors can target important plant functions, while plants evolve to improve effector recognition. Rapid evolution in genes encoding effectors can be facilitated by transposable elements (TEs). In Magnaporthe oryzae, the causal agent of blast disease in several cereals and grasses, TEs play important roles in chromosomal evolution as well as the gain or loss of effector genes in host specialized lineages. However, a global understanding of TE dynamics driving effector evolution at population scale and across lineages is lacking. RESULTS: Here, we focus on 16 AVR effector loci assessed across a global sampling of 11 reference genomes and 447 newly generated draft genome assemblies from publicly available short-read sequencing data across all major M. oryzae lineages and outgroups. We classified each effector based on evidence for duplication, deletion and translocation processes among lineages. Next, we determined AVR gain and loss dynamics across lineages allowing for a broad categorization of effector dynamics. Each AVR was integrated in a distinct genomic niche determined by the TE activity profile contributing to the diversification at the locus. We quantified TE contributions to effector niches and found that TE identity helped diversify AVR loci. We used the large genomic dataset to recapitulate the evolution of the rice blast AVR1-CO39 locus. CONCLUSIONS: Taken together, our work demonstrates how TE dynamics are an integral component of M. oryzae effector evolution, likely facilitating escape from host recognition. In-depth tracking of effector loci is a valuable tool to predict the durability of host resistance.

Ascomycota↗

euHCVdb: the European hepatitis C virus database.

The hepatitis C virus (HCV) genome shows remarkable sequence variability, leading to the classification of at least six major genotypes, numerous subtypes and a myriad of quasispecies within a given host. A database allowing researchers to investigate the genetic and structural variability of all available HCV sequences is an essential tool for studies on the molecular virology and pathogenesis of hepatitis C as well as drug design and vaccine development. We describe here the European Hepatitis C Virus Database (euHCVdb, http://euhcvdb.ibcp.fr), a collection of computer-annotated sequences based on reference genomes. The annotations include genome mapping of sequences, use of recommended nomenclature, subtyping as well as three-dimensional (3D) molecular models of proteins. A WWW interface has been developed to facilitate database searches and the export of data for sequence and structure analyses. As part of an international collaborative effort with the US and Japanese databases, the European HCV Database (euHCVdb) is mainly dedicated to HCV protein sequences, 3D structures and functional analyses.

Databases, Protein↗

A Practical Approach to High-Throughput and Accurate Mapping-by-Sequencing in Arabidopsis.

Forward-directed genetic screens are extremely powerful in identifying novel genes involved in a specific biological process, including various chromatin regulatory pathways. However, the traditional ways of genetic mapping are time- and cost-demanding. Recently, the whole process was revolutionized by the development of mapping-by-sequencing (MBS) protocols. In MBS, the causal mutations and their positions within genes are identified directly by whole-genome sequencing and bioinformatics analysis of the bulk of mutant plants selected based on the mutant phenotype from a segregating population. MBS increases precision and economizes the mapping. Here, we describe a general protocol and provide practical tips on how to proceed with the mapping-by-sequencing on the example of Arabidopsis forward-directed genetic screen designed to identify mutants sensitive to a specific type of DNA damage. The described protocol is generally applicable to a wide range of genetic screens in various inbreeding species with a reference genome sequence.

Arabidopsis↗

'Binding, bending and bonding': polypurine tract-primed initiation of plus-strand DNA synthesis in human immunodeficiency virus.

During the course of reverse transcription, human immunodeficiency virus type 1 reverse transcriptase (HIV-1 RT) initiates plus-strand DNA synthesis from two highly conserved, purine-rich RNA segments of the viral genome referred to as the 3' and central polypurine tracts (3' and cPPTs). Processing of these elements occurs in several sequential steps including (1) minus-strand DNA synthesis over the PPT(s), (2) ribonuclease H (RNase H) mediated cleavage at the PPT 3' terminus, (3) plus-strand DNA synthesis from the nascent RNA primer(s), and (4) primer removal. Completing each of these steps precisely and specifically is essential, as failure to do so can result in reduced virus replication and/or impaired integration of viral DNA into the host cell genome. In this review, plus-strand primer processing in HIV-1 is discussed from biochemical, structural, and historical perspectives. A comparative analysis of PPT-processing in different LTR-containing retroelements is also presented.

DNA, Viral↗

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing↗

Ambient temperature storage of individual parasitic nematode larvae for whole genome sequencing.

Soil-transmitted helminth (STH) infections are a major public health burden, and there are programmes of mass drug administration that attempt to ameliorate the harm that they cause. There has been increasing use of genomics to study STH infections and other parasitic nematodes, with particular interest in whole genome sequencing (WGS). For such studies, samples are commonly stored frozen, but in settings where these infections are endemic this can be difficult, and so there would be advantages to having ambient temperature storage methods. We investigated two ambient temperature storage methods - FTA cards and DESS buffer - for infective larvae of the rat parasites Nippostrongylus brasiliensis and Strongyloides ratti, prior to DNA extraction and then WGS. Our results showed that for individual larvae stored on FTA cards or in DESS buffer, this resulted in a lower proportion of sequence reads that mapped to the reference genomes, compared to the frozen control samples. Generally, for individual larvae, DESS-storage resulted in better sequencing results than FTA-storage. However, for pools of 10 or 50 larvae, then these ambient temperature storage methods generally resulted in comparable sequence read mapping to the frozen control samples.

Animals↗

OCA2 splice site variant in German Spitz dogs with oculocutaneous albinism.

We investigated a German Spitz family where the mating of a black male to a white female had yielded three puppies with an unexpected light brown coat color, lightly pigmented lips and noses, and blue eyes. Combined linkage and homozygosity analysis based on a fully penetrant monogenic autosomal recessive mode of inheritance identified a critical interval of 15 Mb on chromosome 3. We obtained whole genome sequence data from one affected dog, three wolves, and 188 control dogs. Filtering for private variants revealed a single variant with predicted high impact in the critical interval in LOC100855460 (XM_005618224.1:c.377+2T>G LT844587.1:c.-45+2T>G). The variant perfectly co-segregated with the phenotype in the family. We genotyped 181 control dogs with normal pigmentation from diverse breeds including 22 unrelated German Spitz dogs, which were all homozygous wildtype. Comparative sequence analyses revealed that LOC100855460 actually represents the 5'-end of the canine OCA2 gene. The CanFam 3.1 reference genome assembly is incorrect and separates the first two exons from the remaining exons of the OCA2 gene. We amplified a canine OCA2 cDNA fragment by RT-PCR and determined the correct full-length mRNA sequence (LT844587.1). Variants in the OCA2 gene cause oculocutaneous albinism type 2 (OCA2) in humans, pink-eyed dilution in mice, and similar phenotypes in corn snakes, medaka and Mexican cave tetra fish. We therefore conclude that the observed oculocutaneous albinism in German Spitz is most likely caused by the identified variant in the 5'-splice site of the first intron of the canine OCA2 gene.

Albinism, Oculocutaneous↗

Promises and Pitfalls of ctDNA testing in the Management of Cholangiocarcinoma.

Diagnosis and treatment of cholangiocarcinoma is often limited by the availability of tissue biopsies for genomic analysis. Liquid biopsies using blood circulating tumor DNA (ctDNA) have emerged as a valuable and non-invasive alternative to conventional testing. ctDNA analysis has advanced the treatment paradigm for cholangiocarcinoma (CCA) by identifying targetable mutations and molecular mechanisms of treatment resistance. Additionally, it has shown preliminary promise in stratifying patients for adjuvant systemic therapy and enabling earlier detection of relapse. However, current ctDNA platforms face biological and technical challenges that limit their sensitivity for certain mutation types (i.e. gene fusions and amplifications), which are commonly found in CCA. To overcome these hurdles, new sequencing techniques and analytic methods involving artificial intelligence, epigenetic profiling, and diverse reference genomes are being developed. These advanced technologies underscore the promise of ctDNA testing as an indispensable tool in the management and study of CCAs.

Cholangiocarcinoma↗

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals↗

A de novo algorithm for allele reconstruction from Oxford nanopore amplicon reads, with application to CYP2D6.

MOTIVATION: The Oxford Nanopore Technologies' sequencing platform offers a path towards bedside genomics, producing long reads that can completely cover a gene of interest, and detect any known or novel variant the gene contains. However, the analysis of these long reads to identify actionable genotypes remains challenging and typically requires customization depending on the target gene. RESULTS: Here, we describe a generic algorithm to accurately reconstruct allele sequences derived from long-reads of amplicon-based data. Rather than calling variants directly from these long-reads, our method takes a "sequence-first" approach, performing an unbiased reconstruction of the underlying amplicon sequences to generate high-confidence reconstructed allele sequences. This is done without user input of the target gene, allowing for any source amplicon to be reconstructed. These high-confidence reconstructed allele sequences are then compared to the genomic reference sequence of the gene to infer the specific diplotype present in the sample. This approach is agnostic towards the number of genes and alleles present and readily detects novel variants. We demonstrate our approach using three independent data sets for CYP2D6, a diverse and complex gene with over 175 known alleles of clinical significance. We show how our approach can accurately recover validated CYP2D6 diplotypes from 20 Coriell samples covering 14 distinct alleles, using different amplicons, flow cell versions, and depths. This includes inferring occurrences of allele duplication events from relative abundances of each allele, a critical factor for ascribing functional effects to a diplotype. Further, we demonstrate our approach's utility for other genomic regions, including HLA. AVAILABILITY: Custom code is available at the following GitHub repository, along with instructions for use and test data: https://github.com/scottdbrown/allele-reconstruction-long-read-amplicon-data. A snapshot of the code at the time of publication is available on Zenodo.org; doi 10.5281/zenodo.19716004. Raw .fastq sequence data for our three sequencing runs is available at the SRA under Bioproject PRJNA1357883 (https://www.ncbi.nlm.nih.gov/bioproject/1357883).

Alleles↗

Heterogeneous duplications in patients with Pelizaeus-Merzbacher disease suggest a mechanism of coupled homologous and nonhomologous recombination.

We describe genomic structures of 59 X-chromosome segmental duplications that include the proteolipid protein 1 gene (PLP1) in patients with Pelizaeus-Merzbacher disease. We provide the first report of 13 junction sequences, which gives insight into underlying mechanisms. Although proximal breakpoints were highly variable, distal breakpoints tended to cluster around low-copy repeats (LCRs) (50% of distal breakpoints), and each duplication event appeared to be unique (100 kb to 4.6 Mb in size). Sequence analysis of the junctions revealed no large homologous regions between proximal and distal breakpoints. Most junctions had microhomology of 1-6 bases, and one had a 2-base insertion. Boundaries between single-copy and duplicated DNA were identical to the reference genomic sequence in all patients investigated. Taken together, these data suggest that the tandem duplications are formed by a coupled homologous and nonhomologous recombination mechanism. We suggest repair of a double-stranded break (DSB) by one-sided homologous strand invasion of a sister chromatid, followed by DNA synthesis and nonhomologous end joining with the other end of the break. This is in contrast to other genomic disorders that have recurrent rearrangements formed by nonallelic homologous recombination between LCRs. Interspersed repetitive elements (Alu elements, long interspersed nuclear elements, and long terminal repeats) were found at 18 of the 26 breakpoint sequences studied. No specific motif that may predispose to DSBs was revealed, but single or alternating tracts of purines and pyrimidines that may cause secondary structures were common. Analysis of the 2-Mb region susceptible to duplications identified proximal-specific repeats and distal LCRs in addition to the previously reported ones, suggesting that the unique genomic architecture may have a role in nonrecurrent rearrangements by promoting instability.

Base Sequence↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

SUMMARY: Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors." These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac. We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC. Alevin-fry-atac is highly scalable and, when using 32 threads, is 2.8 times faster than Chromap (the second fastest approach) while using only 33% of the memory required by Chromap. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual color-enhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger. AVAILABILITY AND IMPLEMENTATION: Alevin-fry-atac is written in Rust and C++17, and is freely-available under a BSD 3-clause license. It is integrated into piscem (https://github.com/COMBINE-lab/piscem) and alevin-fry (https://github.com/COMBINE-lab/alevin-fry), and is also supported directly as part of simpleaf (https://github.com/COMBINE-lab/simpleaf).

Single-Cell Analysis↗

The L4 22-kilodalton protein plays a role in packaging of the adenovirus genome.

Packaging of the adenovirus (Ad) genome into a capsid is absolutely dependent upon the presence of a cis-acting region located at the left end of the genome referred to as the packaging domain. The functionally significant sequences within this domain consist of at least seven similar repeats, referred to as the A repeats, which have the consensus sequence 5' TTTG-N(8)-CG 3'. In vitro and in vivo binding studies have demonstrated that the adenovirus protein IVa2 binds to the CG motif of the packaging sequences. In conjunction with IVa2, another virus-specific protein binds to the TTTG motifs in vitro. The efficient formation of these protein-DNA complexes in vitro was precisely correlated with efficient packaging activity in vivo. We demonstrate that the binding activity to the TTTG packaging sequence motif is the product of the L4 22-kDa open reading frame. Previously, no function had been ascribed to this protein. Truncation of the L4 22-kDa protein in the context of the viral genome did not reduce viral gene expression or viral DNA replication but eliminated the production of infectious virus. We suggest that the L4 22-kDa protein, in conjunction with IVa2, plays a critical role in the recognition of the packaging domain of the Ad genome that leads to viral DNA encapsidation. The L4 22-kDa protein is also involved in recognition of transcription elements of the Ad major late promoter.

Adenoviridae↗

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses, often with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors…. These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac . We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC . Alevin-fry-atac is highly scalable and, when using 32 threads, is approximately 2.8 times faster than Chromap (the second fastest approach) while using approximately one third of the memory and mapping slightly more reads. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual colorenhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry ) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger . Furthermore, our modified pseudoalignment approach should be easily applicable and extendable to other genome-centric mapping-based tasks and modalities such as standard DNA-seq, DNase-seq, Chip-seq and Hi-C.

Journal Article↗

Sequence variability within the tobacco retrotransposon Tnt1 population.

Retroviruses consist of populations of different but closely related genomes referred to as quasispecies. A high mutation rate coupled with extremely rapid replication cycles allows these sequences to be highly interconnected in a rapid equilibrium. It is not known if other retroelements can show a similar population structure. We show here that when the tobacco Tnt1 retrotransposon is expressed, its RNA is not a unique sequence but a population of different but closely related sequences. Nevertheless, this highly variable population is not in a rapid equilibrium and could not be considered as a quasispecies. We have thus named the structure presented by Tnt1 RNA quasispecies-like. We show that the expression of Tnt1 in different situations gives rise to different populations of Tnt1 RNA sequences, suggesting an adaptive capacity for this element. The analysis of the variability within the total genomic population of Tnt1 elements shows that mutations frequently occur in important regulatory elements and that defective elements are often produced. We discuss the implications that this population structure could have for Tnt1 regulation and evolution.

Base Sequence↗