Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny↗

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article↗

[Annotation of complete genomic sequence of 3p24-p25 478 kb of human DNA].

OBJECTIVE: To annotate the human genome 3p24-p25 478 kb complete sequence. METHODS: The protein-coding genes in the genomic sequence were identified by using ab initio gene finding, homology-based similarity database searching and all or partial mRNA aligning with genomic sequence, and the content feature of the genomic sequence were analyzed by using EMBOSS package. RESULTS: Two known genes SLC6A1 and SLC6A11 were identified; as well as the GC content of this genomic sequence was 47% and 3 putative CpG islands were predicted in the genomic sequence, located in 130,685-131,516 bp, 307,090-307,870 bp and 415,585-416,308 bp, respectively. CONCLUSIONS: The methods, as mentioned above, might be used for annotating the biological information in the genomic sequence, such as gene structure, GC content, CpG island.

Base Sequence↗

Fuzzy Hidden Markov Models: a new approach in multiple sequence alignment.

This paper proposes a novel method for aligning multiple genomic or proteomic sequences using a fuzzyfied Hidden Markov Model (HMM). HMMs are known to provide compelling performance among multiple sequence alignment (MSA) algorithms, yet their stochastic nature does not help them cope with the existing dependence among the sequence elements. Fuzzy HMMs are a novel type of HMMs based on fuzzy sets and fuzzy integrals which generalizes the classical stochastic HMM, by relaxing its independence assumptions. In this paper, the fuzzy HMM model for MSA is mathematically defined. New fuzzy algorithms are described for building and training fuzzy HMMs, as well as for their use in aligning multiple sequences. Fuzzy HMMs can also increase the model capability of aligning multiple sequences mainly in terms of computation time. Modeling the multiple sequence alignment procedure with fuzzy HMMs can yield a robust and time-effective solution that can be widely used in bioinformatics in various applications, such as protein classification, phylogenetic analysis and gene prediction, among others.

Base Sequence↗

Molecular cloning and characterization of the DNAs of human papillomaviruses 19, 20, and 25 from a patient with epidermodysplasia verruciformis.

Five human papillomavirus (HPV) DNAs from lesions of an epidermodysplasia verruciformis patient were cloned in lambda L 47: DNA of HPV 5, which predominated in the carcinoma; DNA of a variant type of HPV 8, which was not detected in the carcinoma DNA by Southern blot hybridization but only by cloning; and DNAs of three papillomaviruses that were isolated from warts. Southern blot and liquid phase DNA-DNA hybridization under stringent conditions showed that the three viruses from warts were new types, which we named HPVs 19, 20, and 25. These viruses cross-hybridized between 3 and 29% among themselves and with HPVs 5 and 8. After physical mapping with several restriction enzymes, the colinear genomes were aligned with HPV 8 DNA to define early and late regions. HPVs 8, 19, and 25 shared homology in different parts of their genomes.

Carcinoma↗

Optimal spliced alignment of homologous cDNA to a genomic DNA template.

MOTIVATION: Supplementary cDNA or EST evidence is often decisive for discriminating between alternative gene predictions derived from computational sequence inspection by any of a number of requisite programs. Without additional experimental effort, this approach must rely on the occurrence of cognate ESTs for the gene under consideration in available, generally incomplete, EST collections for the given species. In some cases, particular exon assignments can be supported by sequence matching even if the cDNA or EST is produced from non-cognate genomic DNA, including different loci of a gene family or homologous loci from different species. However, marginally significant sequence matching alone can also be misleading. We sought to develop an algorithm that would simultaneously score for predicted intrinsic splice site strength and sequence matching between the genomic DNA template and a related cDNA or EST. In this case, weakly predicted splice sites may be chosen for the optimal scoring spliced alignment on the basis of surrounding sequence matching. Strongly predicted splice sites will enter the optimal spliced alignment even without strong sequence matching. RESULTS: We designed a novel algorithm that produces the optimal spliced alignment of a genomic DNA with a cDNA or EST based on scoring for both sequence matching and intrinsic splice site strength. By example, we demonstrate that this combined approach appears to improve gene prediction accuracy compared with current methods that rely only on either search by content and signal or on sequence similarity. AVAILABILITY: The algorithm is available as a C subroutine and is implemented in the SplicePredictor and GeneSeqer programs. The source code is available via anonymous ftp from ftp. zmdb.iastate.edu. Both programs are also implemented as a Web service at http://gremlin1.zool.iastate.edu/cgi-bin/s p.cgiand http://gremlin1.zool.iastate.edu/cgi-bin/g s.cgi, respectively. CONTACT: vbrendel@iastate.edu

Algorithms↗

Lytic properties and genomic analysis of bacteriophage Brt_Psa3, targeting Pseudomonas syringae pv. actinidiae.

Pseudomonas syringae pv. actinidiae (Psa) is the causative agent of bacterial canker in kiwifruit (Actinidia spp.). Psa biovar 3 is the most prevalent and virulent, causing frequent and severe outbreaks worldwide. While current treatments have low efficacy, bacteriophages emerge as possible environmentally safe alternative biocontrol agents. In this study, bacteriophage Brt_Psa3 was isolated from the soil of a kiwifruit orchard in Portugal. Morphologically, Brt_Psa3 forms clear plaques and has a Podoviral morphotype. The bacteriophage exhibited broad lytic activity against several plant-pathogenic Pseudomonas strains, including Psa isolates. The isolated bacteriophage has a latent period of 100 min, a burst size of 143 particles/cell, and demonstrates stability at different temperatures and pH values found in kiwifruit orchards. In addition, Brt_Psa3 exhibited tolerance to UVA irradiation during 120 min of incubation. Brt_Psa3 belongs to the Autographiviridae family and Ghunavirus genus, based on full-genome nucleotide alignment and supported by phylogenetic analysis of structural proteins. The phage contains 51 open reading frames with no antibiotic resistance genes identified, within a genome of 40.509 base pairs. In vitro experiments with kiwifruit leaves demonstrated significant reduction of Psa levels (40%) on leaf surfaces, highlighting the bacteriophage's therapeutic potential in managing bacterial canker in kiwifruits.

Pseudomonas syringae↗

EAnnot: a genome annotation tool using experimental evidence.

The sequence of any genome becomes most useful for biological experimentation when a complete and accurate gene set is available. Gene prediction programs offer an efficient way to generate an automated gene set. Manual annotation, when performed by experienced annotators, is more accurate and complete than automated annotation. However, it is a laborious and expensive process, and by its nature, introduces a degree of variability not found with automated annotation. EAnnot (Electronic Annotation) is a program originally developed for manually annotating the human genome. It combines the latest bioinformatics tools to extract and analyze a wide range of publicly available data in order to achieve fast and reliable automatic gene prediction and annotation. EAnnot builds gene models based on mRNA, EST, and protein alignments to genomic sequence, attaches supporting evidence to the corresponding genes, identifies pseudogenes, and locates poly(A) sites and signals. Here, we compare manual annotation of human chromosome 6 with annotation performed by EAnnot in order to assess the latter's accuracy. EAnnot can readily be applied to manual annotation of other eukaryotic genomes and can be used to rapidly obtain an automated gene set.

Algorithms↗

Significance of interspecies matches when evolutionary rate varies.

We develop techniques to estimate the statistical significance of gap-free alignments between two genomic DNA sequences, using human-mouse alignments as an example. The sequences are assumed to be sufficiently similar that some but not all of the neutrally evolving regions (i.e., those under no evolutionary constraint) can be reliably aligned. Our goal is to model the situation in which the neutral rate of evolution, and hence the extent of the aligning intervals, varies across the genome. In some cases, this permits the weaker of two matches to be judged as less likely to have arisen by chance, provided it lies in a genomic interval with a high level of background divergence. We employ a hidden Markov model to capture variations in divergence rates and assign probability values to gap-free alignments using techniques of Dembo and Karlin, which are related to those used for the same purpose by BLAST. Our methods are illustrated in detail using a 1.49 Mb genomic region. Results obtained from the analysis of human chromosome 22 using these techniques are also provided.

Animals↗

Ancient genomic architecture for mammalian olfactory receptor clusters.

BACKGROUND: Mammalian olfactory receptor (OR) genes reside in numerous genomic clusters of up to several dozen genes. Whole-genome sequence alignment nets of five mammals allow their comprehensive comparison, aimed at reconstructing the ancestral olfactory subgenome. RESULTS: We developed a new and general tool for genome-wide definition of genomic gene clusters conserved in multiple species. Syntenic orthologs, defined as gene pairs showing conservation of both genomic location and coding sequence, were subjected to a graph theory algorithm for discovering CLICs (clusters in conservation). When applied to ORs in five mammals, including the marsupial opossum, more than 90% of the OR genes were found within a framework of 48 multi-species CLICs, invoking a general conservation of gene order and composition. A detailed analysis of individual CLICs revealed multiple differences among species, interpretable through species-specific genomic rearrangements and reflecting complex mammalian evolutionary dynamics. One significant instance involves CLIC #1, which lacks a human member, implying the human-specific deletion of an OR cluster, whose mouse counterpart has been tentatively associated with isovaleric acid odorant detection. CONCLUSION: The identified multi-species CLICs demonstrate that most of the mammalian OR clusters have a common ancestry, preceding the split between marsupials and placental mammals. However, only two of these CLICs were capable of incorporating chicken OR genes, parsimoniously implying that all other CLICs emerged subsequent to the avian-mammalian divergence.

Animals↗

Human retinoblastoma susceptibility gene: genomic organization and analysis of heterozygous intragenic deletion mutants.

A gene in chromosome region 13q14 has been identified as the human retinoblastoma susceptibility (RB) gene on the basis of altered gene expression found in virtually all retinoblastomas. In order to further characterize the RB gene and its structural alterations, we examined genomic clones of the RB gene isolated from both a normal human genomic library and a library made from DNA of the retinoblastoma cell line Y79. First, a restriction and exon map of the RB gene was constructed by aligning overlapping genomic clones, yielding three contiguous regions ("contigs") of 150 kilobases total length separated by two gaps. At least 20 exons were identified in genomic clones, and these were provisionally numbered. Second, two overlapping genomic clones that demonstrated a DNA deletion of exons 2 through 6 from one RB allele were isolated from the Y79 library. To confirm and extend this result, a unique sequence probe from intron 1 was used to detect similar and possibly identical heterozygous deletions in genomic DNA from three retinoblastoma cell lines, thereby explaining the origins of their shortened RB mRNA transcripts. The same probe detected genomic rearrangements in fibroblasts from two hereditary retinoblastoma patients, indicating that intron 1 includes a frequent site for mutations conferring predisposition to retinoblastoma. Third, this probe also detected a polymorphic site for BamHI with allele frequencies near 0.5/0.5. Identification of commonly mutated regions will contribute significantly to genetic diagnosis in retinoblastoma patients and families.

Cell Line↗

In silico analysis of 2085 clones from a normalized rat vestibular periphery 3' cDNA library.

The inserts from 2400 cDNA clones isolated from a normalized Rattus norvegicus vestibular periphery cDNA library were sequenced and characterized. The Wackym-Soares vestibular 3' cDNA library was constructed from the saccular and utricular maculae, the ampullae of all three semicircular canals and Scarpa's ganglia containing the somata of the primary afferent neurons, microdissected from 104 male and female rats. The inserts from 2400 randomly selected clones were sequenced from the 5' end. Each sequence was analyzed using the BLAST algorithm compared to the Genbank nonredundant, rat genome, mouse genome and human genome databases to search for high homology alignments. Of the initial 2400 clones, 315 (13%) were found to be of poor quality and did not yield useful information, and therefore were eliminated from the analysis. Of the remaining 2085 sequences, 918 (44%) were found to represent 758 unique genes having useful annotations that were identified in databases within the public domain or in the published literature; these sequences were designated as known characterized sequences. 1141 sequences (55%) aligned with 1011 unique sequences had no useful annotations and were designated as known but uncharacterized sequences. Of the remaining 26 sequences (1%), 24 aligned with rat genomic sequences, but none matched previously described rat expressed sequence tags or mRNAs. No significant alignment to the rat or human genomic sequences could be found for the remaining 2 sequences. Of the 2085 sequences analyzed, 86% were singletons. The known, characterized sequences were analyzed with the FatiGO online data-mining tool (http://fatigo.bioinfo.cnio.es/) to identify level 5 biological process gene ontology (GO) terms for each alignment and to group alignments with similar or identical GO terms. Numerous genes were identified that have not been previously shown to be expressed in the vestibular system. Further characterization of the novel cDNA sequences may lead to the identification of genes with vestibular-specific functions. Continued analysis of the rat vestibular periphery transcriptome should provide new insights into vestibular function and generate new hypotheses. Physiological studies are necessary to further elucidate the roles of the identified genes and novel sequences in vestibular function.

Afferent Pathways↗

Characterization of the Brassica campestris mitochondrial gene for subunit six of NADH dehydrogenase: nad6 is present in the mitochondrion of a wide range of flowering plants.

We have isolated the Brassica campestris mitochondrial gene nad6, coding for subunit six of NADH dehydrogenase. The deduced amino-acid sequence of this gene shows considerable similarity to mitochondrially encoded NAD6 proteins of other organisms as well as to NAD6 proteins coded for by plant chloroplast DNAs. The B. campestris nad6 gene appears to lack introns and produces an abundant transcript which is comparable in size to a previously described, unidentified transcript (#18) mapped to the B. campestris mitochondrial genome. An alignment of NAD6 proteins (deduced from DNA sequences) suggests that B. campestris nad6 transcripts are edited. Southern-blot hybridization indicates that nad6 is present in the mitochondrial genome of all of a wide range of flowering plant species examined.

Amino Acid Sequence↗

Post-processing long pairwise alignments.

MOTIVATION: The local alignment problem for two sequences requires determining similar regions, one from each sequence, and aligning those regions. For alignments computed by dynamic programming, current approaches for selecting similar regions may have potential flaws. For instance, the criterion of Smith and Waterman can lead to inclusion of an arbitrarily poor internal segment. Other approaches can generate an alignment scoring less than some of its internal segments. RESULTS: We develop an algorithm that decomposes a long alignment into sub-alignments that avoid these potential imperfections. Our algorithm runs in time proportional to the original alignment's length. Practical applications to alignments of genomic DNA sequences are described.

Algorithms↗

Comparative genomics reveals genotype-phenotype concordance and cryptic resistomes in clinical Pseudomonas aeruginosa.

BACKGROUND: Pseudomonas aeruginosa (P. aeruginosa) is a major pathogen because of its adaptability. It shows rapid evolution of multidrug resistance (MDR). Phenotype-based diagnostics often fail to detect silent resistance determinants and early adaptive changes. This study integrates phenotypic profiling with whole-genome sequencing (WGS) to examine resistance architecture in clinical isolates from eastern India. METHODS: From 1295 culture-positive P. aeruginosa specimens collected at a tertiary care hospital in eastern India. Using predefined criteria, representative MDR and non-MDR isolates were selected, including distinct resistance phenotypes, specimen-source diversity, and hospital and community-acquired settings; multivariate analysis of resistance profiles illustrated phenotypic diversity. Antimicrobial susceptibility assessed using VITEK-2 and Kirby-Bauer disk diffusion, species identity confirmed by 16 S rRNA sequencing, and genomic analysis processed through a reference-guided workflow. Antimicrobial Resistance (AMR) determinants were identified through CARD, and phylogenetic tree constructed from 454 publicly available P. aeruginosa genomes. RESULTS: MDR exhibited greater sequence divergence relative to PA14 (~ 69,000 variants) than the non-MDR isolate (~ 58,700 variants), with > 92% coverage at ≥ 30X depth. Strong genotype-phenotype concordance observed in MDR isolates across five antibiotic classes, associated with β-lactamase variants (PDC-67, OXA-396) and regulatory adaptations (ArmR, cprS). The non-MDR isolate harboured gyrA (T83I) resistance-associated mutations, PDC-1, and OXA-847 without phenotypic expression, indicating silent resistome. Phylogenetically, MDR isolates clustered tightly within the phylogeny, while the non-MDR isolate formed a distinct lineage. CONCLUSION: Observed genomic differences align with adaptation under antimicrobial selection, though confirmation requires larger collections. The non-MDR isolate retained a silent resistome. Findings highlight limitations of phenotype-only diagnostics, support genomic data integration, and emphasize transcriptomics for hidden resistance expression and regulatory dynamics.

Pseudomonas aeruginosa↗

Regulatory potential scores from genome-wide three-way alignments of human, mouse, and rat.

We generalize the computation of the Regulatory Potential (RP) score from two-way alignments of human and mouse to three-way alignments of human, mouse, and rat. This requires overcoming technical challenges that arise because the complexity of the models underlying the score increases exponentially with the number of species. Despite the close evolutionary proximity of rat to mouse, we find that adding the rat sequence increases our ability to predict genomic sites that regulate gene transcription. A variant of the RP scoring scheme that accounts for local variation in neutral mutational patterns further improves our predictions.

Actins↗

Synteny between Arabidopsis thaliana and rice at the genome level: a tool to identify conservation in the ongoing rice genome sequencing project.

BLASTX alignment between 189.5 Mb of rice genomic sequence and translated Arabidopsis thaliana annotated coding sequences (CDS) identified 60 syntenic regions involving 4-22 rice orthologs covering < or =3.2 cM (centiMorgan). Most regions are <3 cM in length. A detailed and updated version of a table representing these regions is available on our web site. Thirty-five rice loci match two distinct A.thaliana loci, as expected from the duplicated nature of the A.thaliana genome. One A.thaliana locus matches two distinct rice regions, suggesting that rice chromosomal sequence duplications exist. A high level of rearrangement characterizing the 60 syntenic regions illustrates the ancient nature of the speciation between A.thaliana and rice. The apparent reduced level of microcollinearity implies the dispersion to new genomic locations, via transposon activity, of single or small clusters of genes in the rice genome, which represents a significant additional effector of plant genome evolution.

Arabidopsis↗

Sequencing and comparison of yeast species to identify genes and regulatory elements.

Identifying the functional elements encoded in a genome is one of the principal challenges in modern biology. Comparative genomics should offer a powerful, general approach. Here, we present a comparative analysis of the yeast Saccharomyces cerevisiae based on high-quality draft sequences of three related species (S. paradoxus, S. mikatae and S. bayanus). We first aligned the genomes and characterized their evolution, defining the regions and mechanisms of change. We then developed methods for direct identification of genes and regulatory motifs. The gene analysis yielded a major revision to the yeast gene catalogue, affecting approximately 15% of all genes and reducing the total count by about 500 genes. The motif analysis automatically identified 72 genome-wide elements, including most known regulatory motifs and numerous new motifs. We inferred a putative function for most of these motifs, and provided insights into their combinatorial interactions. The results have implications for genome analysis of diverse organisms, including the human.

Base Sequence↗