Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

e2g: an interactive web-based server for efficiently mapping large EST and cDNA sets to genomic sequences.

e2g is a web-based server which efficiently maps large expressed sequence tag (EST) and cDNA datasets to genomic DNA. It significantly extends the volume of data that can be mapped in reasonable time, and makes this improved efficiency available as a web service. Our server hosts large collections of EST sequences (e.g. 4.1 million mouse ESTs of 1.87 Gb) in precomputed indexed data structures for efficient sequence comparison. The user can upload a genomic DNA sequence of interest and rapidly compare this to the complete collection of ESTs on the server. This delivers a mapping of the ESTs on the genomic DNA. The e2g web interface provides a graphical overview of the mapping. Alignments of the mapped EST regions with parts of the genomic sequence are visualized. Zooming functions allow the user to interactively explore the results. Mapped sequences can be downloaded for further analysis. e2g is available on the Bielefeld University Bioinformatics Server at http://bibiserv.techfak.uni-bielefeld.de/e2g/.

Base Sequence↗

Purification and characterization of the lytic activity induced by the prolate-headed bacteriophage P001 in Lactococcus lactis.

The lytic activity induced by the lactococcal bacteriophage P001 was isolated from phage lysates of Lactococcus lactis by a four-step purification procedure. Two proteins lytic for L. lactis were identified with molecular weights of 28 kDA and 8 kDa, respectively. The N-terminal amino acid sequences of the two proteins were determined and degenerated oligonucleotide probes corresponding to these sequences were synthesized. DNA hybridization experiments with phage P001-DNA and lactococcal DNA revealed that both proteins were apparently encoded by a single lysin gene located on the phage P001 genome. This was confirmed by alignment of the determined N-terminal amino acid sequences with nucleotide sequences which were deduced from cloned Lactococcus bacteriophage lysin genes.

Bacteriophages↗

IS Rm31, a new insertion sequence of the IS 66 family in Sinorhizobium meliloti.

Sinorhizobium meliloti natural populations show a high level of genetic polymorphism possibly due to the presence of mobile genetic elements such as insertion sequences (IS), transposons, and bacterial mobile introns. The analysis of the DNA sequence polymorphism of the nod region of S. meliloti p SymA megaplasmid in an Italian isolate led to the discovery of a new insertion sequence, IS Rm31. IS Rm31 is 2,803 bp long and has 22-bp-long terminal inverted repeat sequences, 8-bp direct repeat sequences generated by transposition, and three ORFs (A, B, C) coding for proteins of 124, 115, and 541 amino acids, respectively. ORF A and ORF C are significantly similar to members of the transposase family. Amino acid and nucleotide sequences indicate that IS Rm31 is a member of the IS 66 family. IS Rm31 sequences were found in 30.5% of the Italian strains analyzed, and were also present in several collection strains of the Rhizobiaceae family, including S. meliloti strain 1021. Alignment of targets sites in the genome of strains carrying IS Rm31 suggested that IS Rm31 inserts randomly into S. meliloti genomes. Moreover, analysis of IS Rm31 insertion sites revealed DNA sequences not present in the recently sequenced S. meliloti strain 1021 genome. In fact, IS Rm31 was in some cases linked to DNA fragments homologous to sequences found in other rhizobia species.

Base Composition↗

Conservation of human alternative splice events in mouse.

Human and mouse genomes share similar long-range sequence organization, and have most of their genes being homologous. As alternative splicing is a frequent and important aspect of gene regulation, it is of interest to assess the level of conservation of alternative splicing. We examined mouse transcript data sets (EST and mRNA) for the presence of transcripts that both make spliced-alignment with the draft mouse genome sequence and demonstrate conservation of human transcript-confirmed alternative and constitutive splice junctions. This revealed 15% of alternative and 67% of constitutive splice junctions as conserved; however, these numbers are patently dependent on the extent of transcript coverage. Transcript coverage of conserved splice patterns is found to correlate well between human and mouse. A model, which extrapolates from observed levels of conservation at increasing levels of transcript support, estimates overall conservation of 61% of alternative and 74% of constitutive splice junctions, albeit with broad confidence intervals. Observed numbers of conserved alternative splicing events agreed with those expected on the basis of the model. Thus, it is apparent that many, and probably most, alternative splicing events are conserved between human and mouse. This, combined with the preservation of alternative frame stop codons in conserved frame breaking events, indicates a high level of commonality in patterns of gene expression between these two species.

Alternative Splicing↗

Genomic organization of an alpha-zein gene cluster in maize.

The genes encoding the alpha-zein proteins of maize constitute a large multigene family of some 75 genes. This multigene family can be divided into four subfamilies based on the nucleotide sequences of their genes and the deduced amino acid sequences of their proteins. We describe for the first time evidence of a clustering of five alpha-zein subfamily 4 (SF4) genes that are members of one of the major alpha-zein subfamilies in a 56 kb region of the genome of the maize inbred line W22. None of the other three known alpha-zein gene subfamilies (SF1, SF2, or SF3) are present in this cluster. The genomic region was reconstructed using restriction endonuclease maps to identify and align three overlapping cosmid clones isolated from a genomic library. The alpha-zein genes are not evenly spaced; the minimum distance between genes is 3.5 kb; the maximum is 13 kb. All the alpha-zein genes in the cluster have the same transcriptional orientation. The location and sequences of some of the repetitive DNA elements in this gene cluster were determined. We estimate that there are a minimum of eight repetitive DNA elements in this region. The sequences of the repetitive elements (not functionally defined) are located between or among the alpha-zein genes. The regions containing two of these repetitive elements (Rep1 and Rep4) have been sequenced; they are about 15 kb apart in the genome. These repetitive elements have similar sequences for about 300 bp out of the 400 bp compared. The regions of sequence similarity, however, are in reverse orientation to one another.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Structural and functional analyses of mutations of the human phenylalanine hydroxylase gene.

BACKGROUND: Phenylketonuria (PKU) is an inborn error of metabolism that results from a deficiency of phenylalanine hydroxylase (PAH). We demonstrated PAH mutational spectrum from patients with PKU, including 10 novel and 3 tetrahydrobiopterin (BH(4))-responsive mutations. In this study, 11 PAH missense mutations, including 6 novel mutations (P69S, G103S, L293M, G332V, S391I, A447P) found in our previous study, 2 mutations common in east Asian patients with PKU (R243Q, R413P), and 3 tetrahydrobiopterin (BH(4))-responsive mutations (R53H, R241C, R408Q) have been functionally and structurally analyzed. METHODS: A transient protein overexpression system and an in vitro BH(4)-responsiveness study were used. The effects of PAH missense mutations on the PAH protein structure were also analyzed. To determine the conservation of 12 mutated residues, PAH was aligned using BLAST against full genomic sequences of 221 different species. Model structures of PAH protein and the composite tetramer were constructed using the software program, SHEBA. RESULTS: No PAH activity was detected for some mutants. However, the residual activities associated with other mutants ranged over a wide spectrum. The missense mutations responsive to BH(4) were not highly conserved throughout the 43 species in the multiple sequence alignment that encode PAH. The composite model structure of PAH revealed that dimer stability was reduced in the BH(4)-responsive mutants, whereas tetramer stability remained normal. CONCLUSION: This expression study analyzed PAH mutations and model structures of mutant PAH proteins are proposed. Correlation between the proposed mutant PAH structures and functions are suggested.

Amino Acid Sequence↗

Evolutionary distance estimation and fidelity of pair wise sequence alignment.

BACKGROUND: Evolutionary distances are a critical measure in comparative genomics and molecular evolutionary biology. A simulation study was used to examine the effect of alignment accuracy of DNA sequences on evolutionary distance estimation. RESULTS: Under the studied conditions, distance estimation was relatively unaffected by alignment error (50% or more of the sites incorrectly aligned) as long as 50% or more of the sites were identical among the sequences (observed P-distance < 0.5). Beyond this threshold, the alignment procedure artificially inflates the apparent sequence identity, skewing distance estimates, and creating alignments that are essentially indistinguishable from random data. This general result was independent of substitution model, sequence length, and insertion and deletion size and rate. CONCLUSION: Examination of the estimated sequence identity may yield some guidance as to the accuracy of the alignment. Inaccurate alignments are expected to have large effects on analyses dependent on site specificity, but analyses that depend on evolutionary distance may be somewhat robust to alignment error as long as fewer than half of the sites have diverged.

Algorithms↗

Validation of the high-throughput marker technology DArT using the model plant Arabidopsis thaliana.

Diversity Arrays Technology (DArT) is a microarray-based DNA marker technique for genome-wide discovery and genotyping of genetic variation. DArT allows simultaneous scoring of hundreds of restriction site based polymorphisms between genotypes and does not require DNA sequence information or site-specific oligonucleotides. This paper demonstrates the potential of DArT for genetic mapping by validating the quality and molecular basis of the markers, using the model plant Arabidopsis thaliana. Restriction fragments from a genomic representation of the ecotype Landsberg erecta (Ler) were amplified by PCR, individualized by cloning and spotted onto glass slides. The arrays were then hybridized with labeled genomic representations of the ecotypes Columbia (Col) and Ler and of individuals from an F(2) population obtained from a Col x Ler cross. The scoring of markers with specialized software was highly reproducible and 107 markers could unambiguously be ordered on a genetic linkage map. The marker order on the genetic linkage map coincided with the order on the DNA sequence map. Sequencing of the Ler markers and alignment with the available Col genome sequence confirmed that the polymorphism in DArT markers is largely a result of restriction site polymorphisms.

Arabidopsis↗

Comparative map alignment of BTA27 and HSA4 and 8 to identify conserved segments of genome containing fat deposition QTL.

Quantitative trait loci (QTL) associated with fat deposition have been identified on bovine Chromosome 27 (BTA27) in two different cattle populations. To generate more informative markers for verification and refinement of these QTL-containing intervals, we initiated construction of a BTA27 comparative map. Fourteen genes were selected for mapping based on previously identified regions of conservation between the cattle and human genomes. Markers were developed from the bovine orthologs of genes found on human Chromosomes 1 (HSA1), 4, 8, and 14. Twelve genes were mapped on the bovine linkage map by using markers associated with single nucleotide polymorphisms or microsatellites. Seven of these genes were also anchored to the physical map by assignment of fluorescence in situ hybridization probes. The remaining two genes not associated with an identifiable polymorphism were assigned only to the physical map. In all, seven genes were mapped to BTA27. Map information generated from the other seven genes not syntenic with BTA27 refined the breakpoint locations of conserved segments between species and revealed three areas of disagreement with the previous comparative map. Consequently, portions of HSA1 and 14 are not conserved on BTA27, and a previously undefined conserved segment corresponding to HSA8p22 was identified near the pericentromeric region of BTA8. These results show that BTA27 contains two conserved segments corresponding to HSA8p, which are separated by a segment corresponding to HSA4q. Comparative map alignment strongly suggests the conserved segment orthologous to HSA8p21-q11 contains QTL for fat deposition in cattle.

Adipose Tissue↗

MUSTANG: a multiple structural alignment algorithm.

Multiple structural alignment is a fundamental problem in structural genomics. In this article, we define a reliable and robust algorithm, MUSTANG (MUltiple STructural AligNment AlGorithm), for the alignment of multiple protein structures. Given a set of protein structures, the program constructs a multiple alignment using the spatial information of the C(alpha) atoms in the set. Broadly based on the progressive pairwise heuristic, this algorithm gains accuracy through novel and effective refinement phases. MUSTANG reports the multiple sequence alignment and the corresponding superposition of structures. Alignments generated by MUSTANG are compared with several handcurated alignments in the literature as well as with the benchmark alignments of 1033 alignment families from the HOMSTRAD database. The performance of MUSTANG was compared with DALI at a pairwise level, and with other multiple structural alignment tools such as POSA, CE-MC, MALECON, and MultiProt. MUSTANG performs comparably to popular pairwise and multiple structural alignment tools for closely related proteins, and performs more reliably than other multiple structural alignment methods on hard data sets containing distantly related proteins or proteins that show conformational changes.

Algorithms↗

Construction of non-symmetric substitution matrices derived from proteomes with biased amino acid distributions.

Automatic comparison of compositionally biased genomes, such as that of the malarial causative agent Plasmodium falciparum (82% adenosine + thymidine), with genomes of average composition, is currently limited. Indeed, popular tools such as BLAST require that amino acid distributions be similar in aligned sequences. However, the P. falciparum genome is so biased that six amino acids account for more than 50% of the protein composition. One reason for the comparison methods failure lies in the compositional difference between the query and the subject proteomes, which is not taken into account in the amino acid substitution matrices. This paper introduces a method to derive substitution matrices, in particular BLOSUM 62, in the frame of the information theory. It allows the construction of non-symmetrical matrices, taking into account the non-symmetric amino acid distributions. The dirAtPf family of matrices allowing the comparison of P. falciparum and A. thaliana is given as an example. This paper further provides an analysis of the obtained matrices in the frame of the information theory, supporting the discrimination advantage they bring.

Amino Acid Sequence↗

Whole genome sequence data on Ethiopian key sorghum landraces and founder lines.

Sorghum (Sorghum bicolor (L.) Moench) is the fifth most important cereal globally. Its genetic diversity is key to improving yield stability, stress tolerance, and adaptation to different environments. Ethiopia is one of the centers for the crop's origin, diversity, and use in both human food and livestock feed. However, genomic data on Ethiopian sorghum remain limited, especially for landraces preferred by local farmers. This dataset consists of whole-genome sequencing data for 188 Ethiopian sorghum accessions, including founder lines and important landraces from major agroecological zones. Sequencing was performed using the Complete Genomics DNBSEQ-T7 platform, generating high-coverage whole-genome data (20 &#xd7; coverage). On average, each accession produced 64.7 million reads. Reads were aligned to the Sorghum bicolor NCBIv3 reference genome and variants called using GATK HaplotypeCaller with joint genotyping (GATK v4.6.1.0). Hard-filtering followed GATK best-practice thresholds (QD <2.0, FS> 60.0, MQ <40.0, MQRankSum <-12.5, ReadPosRankSum <-8.0), retaining biallelic SNPs with mean depth 10-50&#xd7;, missingness &#x2264;20%, and MAF &#x2265;0.05, yielding 6095,752 high-confidence SNPs across 185 accessions. Both the raw FASTQ files and processed VCF files are publicly available to support studies of sorghum genetic diversity, population structure, selection, and the genetic basis of important traits.

Adaptation↗

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence↗

Cyclin A/Cdk1 promotes chromosome alignment and timely mitotic progression.

To ensure genomic fidelity, a series of spatially and temporally coordinated events is executed during prometaphase of mitosis, including bipolar spindle formation, chromosome attachment to spindle microtubules at kinetochores, the correction of erroneous kinetochore-microtubule (k-MT) attachments, and chromosome congression to the spindle equator. Cyclin A/Cdk1 kinase plays a key role in destabilizing k-MT attachments during prometaphase to promote correction of erroneous k-MT attachments. However, it is unknown whether Cyclin A/Cdk1 kinase regulates other events during prometaphase. Here, we investigate additional roles of Cyclin A/Cdk1 in prometaphase by using an siRNA knockdown strategy to deplete endogenous Cyclin A from human cells. We find that depleting Cyclin A significantly extends mitotic duration, specifically prometaphase, because chromosome alignment is delayed. Unaligned chromosomes display erroneous monotelic, syntelic, or lateral k-MT attachments suggesting that bioriented k-MT attachment formation is delayed in the absence of Cyclin A. Mechanistically, chromosome alignment is likely impaired because the localization of the kinetochore proteins BUB1 kinase, KNL1, and MPS1 kinase are reduced in Cyclin A-depleted cells. Moreover, we find that Cyclin A promotes BUB1 kinetochore localization independently of its role in destabilizing k-MT attachments. Thus, Cyclin A/Cdk1 facilitates chromosome alignment during prometaphase to support timely mitotic progression.

Humans↗

Cyclin A/Cdk1 promotes chromosome alignment and timely mitotic progression.

To ensure genomic fidelity a series of spatially and temporally coordinated events are executed during prometaphase of mitosis, including bipolar spindle formation, chromosome attachment to spindle microtubules at kinetochores, the correction of erroneous kinetochore-microtubule (k-MT) attachments, and chromosome congression to the spindle equator. Cyclin A/Cdk1 kinase plays a key role in destabilizing k-MT attachments during prometaphase to promote correction of erroneous k-MT attachments. However, it is unknown if Cyclin A/Cdk1 kinase regulates other events during prometaphase. Here, we investigate additional roles of Cyclin A/Cdk1 in prometaphase by using an siRNA knockdown strategy to deplete endogenous Cyclin A from human cells. We find that depleting Cyclin A significantly extends mitotic duration, specifically prometaphase, because chromosome alignment is delayed. Unaligned chromosomes display erroneous monotelic, syntelic, or lateral k-MT attachments suggesting that bioriented k-MT attachment formation is delayed in the absence of Cyclin A. Mechanistically, chromosome alignment is likely impaired because the localization of the kinetochore proteins BUB1 kinase, KNL1, and MPS1 kinase are reduced in Cyclin A-depleted cells. Moreover, we find that Cyclin A promotes BUB1 kinetochore localization independently of its role in destabilizing k-MT attachments. Thus, Cyclin A/Cdk1 facilitates chromosome alignment during prometaphase to support timely mitotic progression.

Preprint↗

Grass evolution inferred from chromosomal rearrangements and geometrical and statistical features in RNA structure.

The grasses (Poaceae) represent a monophyletic lineage that arose about 70 million years ago. The lineage contains about 10,000 species that differ widely in morphology and physiology. Species show striking differences in genome size, a feature important in the context of conservation of gene content and order (synteny and colinearity) and in the extension of genomic information directly from one grass species to another using comparative approaches. Grass diversification has been a contentious issue, as the exact branching order of the various subfamilies has been difficult to establish with standard methods. This motivated an evolutionary study of deep phylogenetic relationships based on the structure of coding and non-coding RNA molecules and on chromosomal rearrangements. Phylogenetic relationships in the grass family were inferred directly from the structure of RNA using cladistic principles and considerations in statistical mechanics. Coded attributes describing topological and thermodynamic information embedded in RNA molecules were treated as linearly ordered multi-state characters and were polarized by fixing the direction of character transformation toward molecular order. Intrinsically rooted phylogenies derived from the structure of signal recognition particle (SRP) RNA, the mRNA encoded by the early nodulation gene enod40, the small subunit of ribosomal RNA (rRNA), and the internal transcribed spacer ITS1 of rRNA established an order for the diversification of major grass lineages, suggesting a sister relationship of the Pooideae and the PACCAD clade. This same conclusion was reached when large-scale chromosomal rearrangements derived from the comparative genetic mapping of cereal genomes were studied. Chromosomal complements aligned in the most parsimonious manner allowed identification and coding of characters depicting chromosomal translocations, insertions, and linkage block arrangements and the reconstruction of phylogenetic trees based on large-scale chromosomal structure. Congruent reconstruction of deep branching relationships using geometrical and statistical features of RNA structure and orthology and large scale chromosomal recombination events support assumptions of polarization in character argumentation, and fail to falsify the claim that extant grass chromosomes can be considered combinations of linkage blocks of an ancestor of the rice genome. Congruence also suggests that the universal tendency toward order in RNA and the search for the most parsimonious organization of be genome architecture appear to be mutually supported drivers of molecular evolution. The study clarifies the relationship of major clades in the grasses, shows that phylogenetic history can be reconstructed effectively from the combinatorial exchange of chromosomal linkage blocks, and reveals considerable phylogenetic signal embedded in the structure of signal polypeptide-coding mRNA molecules, describing an instance where mRNA structure is the subject of strong evolutionary constraint.

Base Pairing↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗

Ancient differentiation of the H and I haplomes in diploid Hordeum species based on 5S rDNA.

5S rDNA clones from 12 South American diploid Hordeum species containing the HH genome and 3 Eurasian diploid Hordeum species containing the II genome, including the cultivated barley Hordeum vulgare, were sequenced and their sequence diversity was analyzed. The 374 sequenced clones were assigned to "unit classes", which were further assigned to haplomes. Each haplome contained 2 unit classes. The naming of the unit classes reflected the haplomes, viz. both the long H1 and short I1 unit classes were identified with II genome diploids, and both the long H2 and long Y2 unit classes were recognized in South American HH genome diploids. Based upon an alignment of all sequences or alignments of representative sequences, we tested several evolutionary models, and then subjected the parameters of the models to a series of maximum likelihood (ML) analyses and various tests, including the molecular clock, and to a Bayesian evolutionary inference analysis using Markov chain Monte Carlo (MCMC). The best fitting model of nucleotide substitution was the HKY+G (Hasegawa, Kishino, Yano 1985 model with the Gamma distribution rates of nucleotide substitutions). Results from both ML and MCMC imply that the long H1 and short I unit classes found in the II genome diploids diverged from each other at the same rate as the long H2 and long Y2 unit classes found in the HH genome diploids. The divergence among the unit classes, estimated to be circa 7 million years, suggests that the genus Hordeum may be a paleopolyploid.

DNA, Plant↗