Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

The Arabidopsis genome: a foundation for plant research.

The sequence of the first plant genome was completed and published at the end of 2000. This spawned a series of large-scale projects aimed at discovering the functions of the 25,000+ genes identified in Arabidopsis thaliana (Arabidopsis). This review summarizes progress made in the past five years and speculates about future developments in Arabidopsis research and its implications for crop science. The provision of large populations of gene disruption lines to the research community has greatly accelerated the impact of genomics on many areas of plant science. The tools and community organization required for plant integrative and systems biology approaches are now ready to accomplish the next big step in plant biology--the integration of knowledge and modeling of biological processes. In the future, plant science will continue to be enriched by the alignment of high-quality basic research (generally conducted in Arabidopsis), with strategic objectives in crop plants. The sequence and analysis of an increasing number of crop plant genomes enhance this alignment and provide new insights into genome evolution and crop plant domestication.

Arabidopsis↗

YASS: enhancing the sensitivity of DNA similarity search.

YASS is a DNA local alignment tool based on an efficient and sensitive filtering algorithm. It applies transition-constrained seeds to specify the most probable conserved motifs between homologous sequences, combined with a flexible hit criterion used to identify groups of seeds that are likely to exhibit significant alignments. A web interface (http://www.loria.fr/projects/YASS/) is available to upload input sequences in fasta format, query the program and visualize the results obtained in several forms (dot-plot, tabular output and others). A standalone version is available for download from the web page.

Algorithms↗

Genomic resources for comparative analyses of obligate avian brood parasitism.

Examples of convergent evolution, wherein distantly related organisms evolve similar traits, including behaviors, underscore the adaptive power of natural selection. In birds, obligate brood parasitism, and the associated loss of parental care behaviors, has evolved independently in seven different lineages, though little is known about the genetic basis of the complex suite of traits associated with this rare life history strategy. We generated genome assemblies for ten brood parasitic species plus eight species representatives of their parental/nesting outgroups. This includes nine long-read chromosome-level assemblies, with scaffold N50 sizes ranging from 38.1 to 72.6 MB, and gene representation completeness measures >97%. Leveraging this new catalog of avian genomes, we constructed clade-level alignments that reveal variation in chromosomal synteny, provide first-time or improved annotations of protein-coding and non-coding genes, and define cross-species ortholog reference sets. We also refine estimates for the timing of the seven independent origins of brood parasitism, ranging from recent events such as 1.6 to 4.5 million years ago in Molothrus cowbirds to much earlier origins over 30 million years ago in two of the three cuckoo lineages. These genomic resources lay the foundation for investigating the genetic and genomic underpinnings of brood parasitism, including the loss of parental care, shifts in mating systems, perhaps resulting in heightened sperm competition, elevated annual fecundity, improved spatial cognition related to nest-finding, and the diverse adaptations shaped by intense coevolution with host species.

assemblies↗

Gene repair using chimeric RNA/DNA oligonucleotides.

An experimental strategy has been developed for the site-specific alteration of genomic DNA. The approach is based on the observation that oligonucleotides containing complementary RNA/DNA hybrid regions are more active than duplex DNA in homologous pairing reactions in vitro. The chimeric molecules are designed with a homologous targeting sequence comprised of a DNA region flanked by blocks of 2'-O-methyl RNA residues (the chimeric strand), its complementary all-DNA strand, thymidine hairpin caps, a single-strand break, and a double-stranded clamp region. The oligonucleotide can align in perfect register with a genomic target except for the designed single base pair mismatch, which is recognized and corrected by harnessing the cell's endogenous DNA repair system. The mechanism of repair has been studied using mammalian cell-free extracts and bacterial systems and has revealed a mismatch correction pathway distinct from homologous recombination. The chimeric molecules have been demonstrated to be effective in the alteration of single nucleotides in episomal and genomic DNA in cell culture, as well as genomic DNA of cells in situ. This is a potentially powerful strategy for gene repair for the myriad hepatic genetic diseases caused by point mutations.

Animals↗

Fast and cheap genome wide haplotype construction via optical mapping.

We describe an efficient algorithm to construct genome wide haplotype restriction maps of an individual by aligning single molecule DNA fragments collected with Optical Mapping technology. Using this algorithm and small amount of genomic material, we can construct the parental haplotypes for each diploid chromosome for any individual. Since such haplotype maps reveal the polymorphisms due to single nucleotide differences (SNPs) and small insertions and deletions (RFLPs), they are useful in association studies, studies involving genomic instabilities in cancer, and genetics, and yet incur relatively low cost and provide high throughput. If the underlying problem is formulated as a combinatorial optimization problem, it can be shown to be NP-complete (a special case of K-population problem). But by effectively exploiting the structure of the underlying error processes and using a novel analog of the Baum-Welch algorithm for HMM models, we devise a probabilistic algorithm with a time complexity that is linear in the number of markers for an epsilon-approximate solution. The algorithms were tested by constructing the first genome wide haplotype restriction map of the microbe T. pseudoana, as well as constructing a haplotype restriction map of a 120 Mb region of Human chromosome 4. The frequency of false positives and false negatives was estimated using simulated data. The empirical results were found very promising.

Algorithms↗

Physical map of the Azoarcus sp. strain BH72 genome based on a bacterial artificial chromosome library as a platform for genome sequencing and functional analysis.

Azoarcus sp. strain BH72 is a Gram-negative proteobacterium of the beta subclass; it is a diazotrophic endophyte of graminaceous plants and can provide significant amounts of fixed nitrogen to its host plant Kallar grass. We aimed to obtain a physical map of the Azoarcus sp. strain BH72 chromosome to be directly used in functional analysis and as a part of an Azoarcus sp. BH72 genome project. A bacterial artificial chromosome (BAC) library was constructed and analysed. A representative physical map with a high density of marker genes was developed in which 64 aligned BAC clones covered almost the entire genome.

Azoarcus↗

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n = 53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans↗

Sequence evidence for RNA recombination in field isolates of avian coronavirus infectious bronchitis virus.

Under laboratory conditions coronaviruses were shown to have a high frequency of recombination. In The Netherlands, vaccination against infectious bronchitis virus (IBV) is performed with vaccines that contain several life-attenuated virus strains. These highly effective vaccines may create ideal conditions for recombination, and could therefore be dangerous in the long term. This paper addresses the question of the frequency of recombination of avian coronavirus IBV in the field. A method was sought to detect and quantify recombination from sequence data. Nucleotide sequences of eight IBV isolates in a region of the genome suspected to contain recombination, were aligned and compared. Phylogenetic trees were constructed for different sections of this region. Differences in topology between these trees were observed, suggesting that in three out of eight strains in vivo RNA recombinant had occurred.

Amino Acid Sequence↗

Functional insights from structural predictions: analysis of the Escherichia coli genome.

Fold assignments for proteins from the Escherichia coli genome are carried out using BASIC, a profile-profile alignment algorithm, recently tested on fold recognition benchmarks and on the Mycoplasma genitalium genome and PSI BLAST, the newest generation of the de facto standard in homology search algorithms. The fold assignments are followed by automated modeling and the resulting three-dimensional models are analyzed for possible function prediction. Close to 30% of the proteins encoded in the E. coli genome can be recognized as homologous to a protein family with known structure. Most of these homologies (23% of the entire genome) can be recognized both by PSI BLAST and BASIC algorithms, but the latter recognizes an additional 260 homologies. Previous estimates suggested that only 10-15% of E. coli proteins can be characterized this way. This dramatic increase in the number of recognized homologies between E. coli proteins and structurally characterized protein families is partly due to the rapid increase of the database of known protein structures, but mostly it is due to the significant improvement in prediction algorithms. Knowing protein structure adds a new dimension to our understanding of its function and the predictions presented here can be used to predict function for uncharacterized proteins. Several examples, analyzed in more detail in this paper, include the DPS protein protecting DNA from oxidative damage (predicted to be homologous to ferritin with iron ion acting as a reducing agent) and the ahpC/tsa family of proteins, which provides resistance to various oxidating agents (predicted to be homologous to glutathione peroxidase).

Algorithms↗

Modeling DNA base substitution in large genomic regions from two organisms.

We studied the substitution patterns in 7661 well-conserved human-mouse alignments corresponding to the intergenic regions of human chromosome 22. Alignments with a high average GC content tend to have a higher human GC content than mouse GC content, indicating a lack of stationarity. Segmenting the alignments into four groups of GC content and fitting the general reversible substitution model (REV) separately gave significantly better fits than the overall fit and the levels of fit are close to that expected under an REV model. In addition, most of the fitted rate matrices are not of the HKY type but are remarkably strand-symmetric, and we constructed a number of substitution matrices that should be useful for genomic DNA sequence alignment. We did not find obvious signs of temporal inhomogeneity in the substitution rates and concluded that the conserved intergenic regions in human chromosome 22 and mouse appear to have evolved from their common ancestors via a process that is approximately reversible and strand-symmetric, assuming site homogeneity and independence.

Animals↗

Imprinted chromosomal regions of the human genome display sex-specific meiotic recombination frequencies.

BACKGROUND: Meiotic recombination events do not occur randomly along a chromosome, but appear to be restricted to specific regions. In addition, some regions in the genome undergo recombination more frequently in the germ cells of one sex than the other. Genomic imprinting, the process by which the two parental alleles of a gene are differentially marked, is another genetic phenomenon associated with inheritance from only one parent or the other. The mechanisms that control meiotic recombination and genomic imprinting are unknown, but both phenomena necessarily depend on the presence of some DNA signal sequences and/or on the structure of the surrounding chromatin domain. RESULTS: In the present study, we compared the frequencies of sex-specific recombination events in three chromosomal regions of the human genome that contain clustered imprinted genes. Alignment of the genetic and physical maps of the ZNF127-SNRPN-IPW-PAR-5-PAR-1 region on chromosome 15q11-q13 (associated with Prader-Willi and Angelman syndromes) and the IGF2-H19 region on chromosome 11p15.5 (associated with Beckwith-Wiedemann syndrome) shows that both regions recombine with very high frequency during male meiosis, and with very low frequency during female meiosis. A third region around the WT-1 gene on chromosome 11p13 also recombines with higher frequency during male meiosis. CONCLUSIONS: The results show that the two best-known imprinted regions in the human genome are characterized by significant differences in recombination frequency during male and female meioses. A third, less well-characterized, imprinted region shows a similar sex-specific bias. On the basis of these observations, we propose a model suggesting that the region-specific differential accessibility of DNA that leads to differential recombination rates during male and female meioses also leads to the male- and female-specific modification of the signal sequences that control genomic imprinting.

Base Sequence↗

BTW: a web server for Boltzmann time warping of gene expression time series.

UNLABELLED: Dynamic time warping (DTW) is a well-known quadratic time algorithm to determine the smallest distance and optimal alignment between two numerical sequences, possibly of different length. Originally developed for speech recognition, this method has been used in data mining, medicine and bioinformatics. For gene expression time series data, time warping distance is arguably a more flexible tool to determine genes having similar temporal expression, hence possibly related biological function, than either Euclidean distance or correlation coefficient--especially since time warping accommodates sequences of different length. The BTW web server allows a user to upload two tab-separated text files A,B of gene expression data, each possibly having a different number of time intervals of different durations. BTW then computes time warping distance between each gene of A with each gene of B, using a recently developed symmetric algorithm which additionally computes the Boltzmann partition function and outputs Boltzmann pair probabilities. The Boltzmann pair probabilities, not available with any other existent software, suggest possible biological significance of certain positions in an optimal time warping alignment. AVAILABILITY: http://bioinformatics.bc.edu/clotelab/BTW/.

Algorithms↗

Genome-wide analysis of SPAK/OSR1 binding motifs.

Based on the alignment of 12 sequences of protein motifs that interact with the kinases SPAK (Ste20-related proline alanine-rich kinase) and OSR1 (oxidative stress response 1), we performed genome-wide searches of the sequence [S/G/V]RFx[V/I]xx[V/I/T/S]xx, where x represents any amino acid. The "Mus musculus" search resulted in the identification of 131 mouse proteins containing 137 SPAK/OSR1 putative binding motifs. Similar numbers were found for human, zebrafish, fruit fly, and worm. A little more than half of the mouse proteins containing SPAK/OSR1 binding domains (53%) were also identified in the human search, whereas approximately 17-18% of these common hits were identified in the zebrafish search. The mouse proteins could be divided into two broad categories: 2/3 had an identified function, whereas 1/3 were either predicted or of unknown function. The known proteins were grouped as transport proteins, other membrane proteins, kinases, phosphatases, cytoskeletal, ribosomal, nuclear, enzymes, and others. Analysis of the location of the SPAK/OSR1 binding motif within the protein sequence revealed distribution throughout the entire length, but with preference to the extreme amino- or carboxyl termini for a large number of proteins. Analysis of the amino acid composition of the motifs revealed a preponderance of serine residues at positions 5, 6, 7, and 8. In summary, our new search found and thus confirms the 12 proteins previously shown to interact with the kinases and identifies 119 potential new targets for SPAK and OSR1 in the mouse proteome.

Amino Acid Motifs↗

Identification and analysis of the promoter region of the human hyaluronan synthase 2 gene.

Hyaluronan (HA) is a linear glycosaminoglycan of the vertebrate extracellular matrix that is synthesized at the plasma membrane by the HA synthase (HAS) enzymes HAS1, -2 and -3. The regulation of HA synthesis has been implicated in a variety of extracellular matrix-mediated and pathological processes, including renal fibrosis. We have recently described the genomic structures of each of the human HAS genes. In the present study, we analyzed the HAS2 promoter region. In 5'-rapid amplification of cDNA ends analysis of purified mRNA from human renal epithelial proximal tubular cells, we detected an extended sequence for HAS2 exon 1, relocating the transcription initiation site 130 nucleotides upstream of the reference HAS2 mRNA sequence, GenBank accession number NM_005328. A luciferase reporter gene assay of nested fragments spanning the 5' terminus of NM_005328 demonstrated the constitutive promoter activity of sequences directly upstream of the repositioned transcription initiation site but not of the newly designated exonic nucleotides. Using reverse transcription-PCR, expression of this extended HAS2 mRNA was demonstrated in a variety of human cell types, and orthologous sequences were detected in mouse and rat kidney. Alignment of human, murine, and equine genomic DNA sequences upstream of the repositioned HAS2 exon 1 provided evidence for the evolutionary conservation of specific transcription factor binding sites. The location of the HAS2 promoter will facilitate analysis of the transcriptional regulation of this gene in a variety of pathological contexts as well as in developmental models in which HAS2 null animals have an embryonic lethal phenotype.

Animals↗

RRNPP quorum-sensing repertoires in the salivarius group genomes: overrepresentation and synchronous activation of SHP/Rgg systems in Streptococcus thermophilus.

UNLABELLED: In Bacillota, quorum sensing can be mediated by RRNPP regulators that are activated by autoinducing peptides (AIPs). In this study, we derived a hidden Markov model profile from a 3D-informed alignment to establish RRNPP repertoires for 527 genomes of streptococci in the salivarius group and identified probable AIPs. The salivarius group encompasses Streptococcus salivarius and Streptococcus vestibularis, which are part of the normal human oral microflora, and Streptococcus thermophilus, one of the most widely used bacteria in the dairy industry. We observed a large amount of plasticity in these repertoires, as well as profound differences among species. Notably, S. salivarius displayed an accumulation of ComR regulators, while S. thermophilus displayed an accumulation of Rgg regulators. The latter family included SHP-associated Rgg regulators, systems in which SHPs serve as AIPs; most of these regulators control the production of post-translationally modified peptides (RaS-RiPPs). Their level of richness contrasts with the genome reduction that accompanied S. thermophilus' adaptation to milk. We then used liquid chromatography-high resolution tandem mass spectrometry to analyze the activity of the eight most common SHP/Rgg systems by characterizing the SHPs and RaS-RiPPs found in the supernatants. We detected four SHPs and one RaS-RiPP that have never been seen before in S. thermophilus, and we showed that seven of the eight SHP/Rgg systems were functional. Finally, by simultaneously monitoring the amounts of both the SHPs and RaS-RiPPs, we demonstrated that the fates of these two peptide types differed during growth. SHP presence in the supernatant was transient, a pattern likely related to the peptides' signaling role. IMPORTANCE: Streptococcus thermophilus possesses an unusually high number of Rgg regulators, which are activated by SHP pheromones that control the production of RaS-RiPPs, peptides with cyclization motifs and growth inhibition properties. We conducted an in silico analysis of regulator repertoires across a wide range of strains; a subsequent experimental study revealed that the majority of the SHP/Rgg systems were functional. Employing an optimized liquid chromatography-high resolution tandem mass spectrometry protocol, we were able to better detect and follow SHP and RaS-RiPP accumulation. While RaS-RiPPs accumulated during growth, SHPs were only transiently present in the extracellular environment. This observation suggests that we could manipulate quorum sensing by adding SHPs to the growth medium and highlights the need to study the functions of the RaS-RiPPs.

Streptococcus thermophilus↗

Tropheryma whipplei Twist: a human pathogenic Actinobacteria with a reduced genome.

The human pathogen Tropheryma whipplei is the only known reduced genome species (<1 Mb) within the Actinobacteria [high G+C Gram-positive bacteria]. We present the sequence of the 927303-bp circular genome of T. whipplei Twist strain, encoding 808 predicted protein-coding genes. Specific genome features include deficiencies in amino acid metabolisms, the lack of clear thioredoxin and thioredoxin reductase homologs, and a mutation in DNA gyrase predicting a resistance to quinolone antibiotics. Moreover, the alignment of the two available T. whipplei genome sequences (Twist vs. TW08/27) revealed a large chromosomal inversion the extremities of which are located within two paralogous genes. These genes belong to a large cell-surface protein family defined by the presence of a common repeat highly conserved at the nucleotide level. The repeats appear to trigger frequent genome rearrangements in T. whipplei, potentially resulting in the expression of different subsets of cell surface proteins. This might represent a new mechanism for evading host defenses. The T. whipplei genome sequence was also compared to other reduced bacterial genomes to examine the generality of previously detected features. The analysis of the genome sequence of this previously largely unknown human pathogen is now guiding the development of molecular diagnostic tools and more convenient culture conditions.

Actinomycetales↗

Physical map of the genome of Vibrio cholerae 569B and localization of genetic markers.

A combined physical and genetic map of the genome of the classical O1 hypertoxinogenic strain 569B of Vibrio cholerae has been constructed. The enzymes NotI, SfiI and CeuI generated DNA fragments of suitable size distribution that could be resolved by pulsed-field gel electrophoresis. The digests produced 37, 22, and 7 fragments, respectively. The CeuI maps of the genomes of strains 569B and O395, constructed by partial restriction digestion, were identical, and the data are consistent with the concept of circular chromosomes. The genome size of each of the strains was estimated to be about 3.2 Mb. The NotI and SfiI digestion profiles of the genomic DNAs of strains 569B and O395 exhibited distinct restriction fragment length polymorphism. The linkages between the 37 NotI fragments of the genome of strain 569B were determined by combining three approaches: isolation of linking clones, analysis of partial digestion fragments, and identification of NotI fragments in isolated CeuI and SfiI fragments. To align linked fragments precisely, NotI-digested genomic DNA was end labeled and separated in the same gel with the NotI-digested DNA to be probed with linking clones. This also allowed the identification of smaller restriction fragments that are not visible in ethidium bromide-stained gels. The presence of repetitive DNA sequences in the V. cholerae 569B genome has been demonstrated. Twenty cloned homologous and heterologous genes and seven rrn operons have been positioned on the physical map. The two copies of the Ctx genetic element in the genome of strain 569B are located about 1,000 kb apart.

Base Sequence↗