Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Cloning and characterization of a human delayed rectifier potassium channel gene.

A human genomic DNA library was screened for sequences homologues to the rat delayed rectifier Kv 2.1 (DRK1) K+ channel cDNA. Three phages were isolated which hybridized to Kv 2.1 cDNA probes. Alignment of the human genomic DNA sequence with the rat cDNA sequence indicated that the open reading frame (ORF) is interrupted by a large intervening sequence, that separates exons encoding the membrane spanning core region of the K+ channel polypeptide. The Kv 2.1 gene occurs once in the human genome and has been mapped to chromosome 20. The human, mouse and rat Kv 2.1 proteins have been highly conserved, showing only a few substitutions outside of the membrane spanning domains in the amino- and carboxy-terminal cytoplasmic domains. Nevertheless, expression of human DRK1 channels in Xenopus oocytes showed that mouse, rat and human Kv 2.1 channels have distinct pharmacological and electrophysiological properties. The observed differences in activation, voltage-dependence, 4-aminopyridine sensitivity and single-channel conductance have to be attributed to amino acid substitutions in the amino-and/or carboxy-terminal cytoplasmic domains. Obviously, these domains of Kv 2.1 channels influence biophysical K+ channel properties, which are thought to be determined solely by the membrane spanning core domain of potassium channels.

Amino Acid Sequence↗

Quantifying ascertainment bias and species-specific length differences in human and chimpanzee microsatellites using genome sequences.

Surveys of variability of homologous microsatellite loci among species reveal an ascertainment bias for microsatellite length where microsatellite loci isolated in one species tend to be longer than homologous loci in related species. Here, we take advantage of the availability of aligned human and chimpanzee genome sequences to compare length difference of homologous microsatellites for loci identified in humans to length difference for loci identified in chimpanzees. We are able to quantify ascertainment bias for a range of motifs and microsatellite lengths. Because ascertainment bias should not exist if a microsatellite selected in one species is as likely to be longer as it is to be shorter than its homologue, we propose that the nature of ascertainment bias can provide evidence for understanding how microsatellites evolve. We show that bias is greater for longer microsatellites but also that many long microsatellites have short homologues. These results are consistent with the notion that growth of long microsatellites is constrained by an upper length boundary that, when reached, sometimes results in large deletions. By evaluating ascertainment bias separately for interrupted and uninterrupted repeats we also show that long microsatellites tend to become interrupted, thereby contributing a second component of ascertainment bias. Having accounted for ascertainment bias, in agreement with results published elsewhere, we find that microsatellites in humans are longer on average than those in chimpanzees. This length difference is similar among repeat motifs but surprisingly comprises two roughly equal components, one associated with the repeats themselves and one with the flanking sequences. The differences we find can only be explained if microsatellites are both evolving directionally under a biased mutation process and are doing so at different rates in different closely related species.

Animals↗

Detection of non-coding RNAs on the basis of predicted secondary structure formation free energy change.

BACKGROUND: Non-coding RNAs (ncRNAs) have a multitude of roles in the cell, many of which remain to be discovered. However, it is difficult to detect novel ncRNAs in biochemical screens. To advance biological knowledge, computational methods that can accurately detect ncRNAs in sequenced genomes are therefore desirable. The increasing number of genomic sequences provides a rich dataset for computational comparative sequence analysis and detection of novel ncRNAs. RESULTS: Here, Dynalign, a program for predicting secondary structures common to two RNA sequences on the basis of minimizing folding free energy change, is utilized as a computational ncRNA detection tool. The Dynalign-computed optimal total free energy change, which scores the structural alignment and the free energy change of folding into a common structure for two RNA sequences, is shown to be an effective measure for distinguishing ncRNA from randomized sequences. To make the classification as a ncRNA, the total free energy change of an input sequence pair can either be compared with the total free energy changes of a set of control sequence pairs, or be used in combination with sequence length and nucleotide frequencies as input to a classification support vector machine. The latter method is much faster, but slightly less sensitive at a given specificity. Additionally, the classification support vector machine method is shown to be sensitive and specific on genomic ncRNA screens of two different Escherichia coli and Salmonella typhi genome alignments, in which many ncRNAs are known. The Dynalign computational experiments are also compared with two other ncRNA detection programs, RNAz and QRNA. CONCLUSION: The Dynalign-based support vector machine method is more sensitive for known ncRNAs in the test genomic screens than RNAz and QRNA. Additionally, both Dynalign-based methods are more sensitive than RNAz and QRNA at low sequence pair identities. Dynalign can be used as a comparable or more accurate tool than RNAz or QRNA in genomic screens, especially for low-identity regions. Dynalign provides a method for discovering ncRNAs in sequenced genomes that other methods may not identify. Significant improvements in Dynalign runtime have also been achieved.

Algorithms↗

B19 virus genome diversity: epidemiological and clinical correlations.

Genetic analysis of parvovirus B19 has been carried out mainly to establish a framework to track molecular epidemiology of the virus and to correlate sequence variability with different pathological and clinical manifestations of the virus. A good amount of information regarding B19 virus sequence variability is available, and presently there are about 400 sequence records deposited in the nucleotide database of NCBI. A few are almost complete genomic sequences, and these allow the construction of a global alignment framework. Many others are partial genomic sequences, limited to selected regions, and these allow comparison of a higher number of isolates from well-defined epidemiological settings and/or pathological conditions. Most studies showed that the genetic variability of B19 virus is low, that molecular epidemiology is possible only on a limited geographical and temporal setting, and that no clear correlations are present between genome sequence and distinctive pathological and clinical manifestations. More recently, several viral isolates have been identified that show remarkable sequence diversity with respect to reference sequences. The identification of variant isolates added to the knowledge of genetic diversity in this virus group and allowed the identification of three divergent genetic clusters, about 10% divergent from each other and still quite distinct from other parvoviruses, that can be thought of as different genotypes within the human erythrovirus group and that show clearly resolved phylogenetic relationship. These variant isolates pose interesting questions regarding the real extent of genetic variability in the human erythroviruses, the relevance of these viruses in terms of epidemiology and their possible implication in the pathogenesis of erythrovirus-related diseases.

Genes, Viral↗

Improved spliced alignment from an information theoretic approach.

MOTIVATION: mRNA sequences and expressed sequence tags represent some of the most abundant experimental data for identifying genes and alternatively spliced products in metazoans. These transcript sequences are frequently studied by aligning them to a genomic sequence template. For existing programs, error-prone, polymorphic and cross-species data, as well as non-canonical splice sites, still present significant barriers to producing accurate, complete alignments. RESULTS: We took a novel approach to spliced alignment that meaningfully combined information from sequence similarity with that obtained from PSSM splice site models. Scoring systems were chosen to maximize their power of discrimination, and dynamic programming (DP) was employed to guarantee optimal solutions would be found. The resultant program, EXALIN, performed better than other popular tools tested under a wide range of conditions that included detection of micro-exons and human-mouse cross-species comparisons. For improved speed with only a marginal decrease in splice site prediction accuracy, EXALIN could perform limited DP guided by a result from BLASTN. AVAILABILITY: The source code, binaries, scripts, scoring matrices and splice site models for human, mouse, rice and Caenorhabditis elegans utilized in this study are posted at http://blast.wustl.edu/exalin. The software (scripts, source code and binaries) is copyrighted but free for all to use.

Algorithms↗

The complete nucleotide sequence and RNA editing content of the mitochondrial genome of rapeseed (Brassica napus L.): comparative analysis of the mitochondrial genomes of rapeseed and Arabidopsis thaliana.

The entire mitochondrial genome of rapeseed (Brassica napus L.) was sequenced and compared with that of Arabidopsis thaliana. The 221 853 bp genome contains 34 protein-coding genes, three rRNA genes and 17 tRNA genes. This gene content is almost identical to that of Arabidopsis: However the rps14 gene, which is a pseudo-gene in Arabidopsis, is intact in rapeseed. On the other hand, five tRNA genes are missing in rapeseed compared to Arabidopsis, although the set of mitochondrially encoded tRNA species is identical in the two Cruciferae. RNA editing events were systematically investigated on the basis of the sequence of the rapeseed mitochondrial genome. A total of 427 C to U conversions were identified in ORFs, which is nearly identical to the number in Arabidopsis (441 sites). The gene sequences and intron structures are mostly conserved (more than 99% similarity for protein-coding regions); however, only 358 editing sites (83% of total editings) are shared by rapeseed and Arabidopsis: Non-coding regions are mostly divergent between the two plants. One-third (about 78.7 kb) and two-thirds (about 223.8 kb) of the rapeseed and Arabidopsis mitochondrial genomes, respectively, cannot be aligned with each other and most of these regions do not show any homology to sequences registered in the DNA databases. The results of the comparative analysis between the rapeseed and Arabidopsis mitochondrial genomes suggest that higher plant mitochondria are extremely conservative with respect to coding sequences and somewhat conservative with respect to RNA editing, but that non-coding parts of plant mitochondrial DNA are extraordinarily dynamic with respect to structural changes, sequence acquisition and/or sequence loss.

Amino Acid Sequence↗

A probabilistic model of 3' end formation in Caenorhabditis elegans.

The 3' ends of mRNAs terminate with a poly(A) tail. This post-transcriptional modification is directed by sequence features present in the 3'-untranslated region (3'-UTR). We have undertaken a computational analysis of 3' end formation in Caenorhabditis elegans. By aligning cDNAs that diverge from genomic sequence at the poly(A) tract, we accurately identified a large set of true cleavage sites. When there are many transcripts aligned to a particular locus, local variation of the cleavage site over a span of a few bases is frequently observed. We find that in addition to the well-known AAUAAA motif there are several regions with distinct nucleotide compositional biases. We propose a generalized hidden Markov model that describes sequence features in C.elegans 3'-UTRs. We find that a computer program employing this model accurately predicts experimentally observed 3' ends even when there are multiple AAUAAA motifs and multiple cleavage sites. We have made available a complete set of polyadenylation site predictions for the C.elegans genome, including a subset of 6570 supported by aligned transcripts.

3' Untranslated Regions↗

Prokaryote phylogeny without sequence alignment: from avoidance signature to composition distance.

This is a review of a new and essentially simple method of inferring phylogenetic relationships from complete genome data without using sequence alignment. The method is based on counting the appearance frequency of oligopeptides of a fixed length (up to K = 6) in the collection of protein sequences of a species. It is a method without fine adjustment and choice of genes. Applied to prokaryotic genomes it has led to results comparable with the bacteriologists' systematics as reflected in the latest 2002 outline of the Bergey's Manual of Systematic Bacteriology. The method has also been used to compare chloroplast genomes and to the phylogeny of Coronaviruses including human SARS-CoV. A key point in our approach is subtraction of a random background from the original counts by using a Markov model of order K-2 in order to highlight the shaping role of natural selection. The implications of the subtraction procedure is specially analyzed and further development of the new approach is indicated.

Algorithms↗

The nucleotype, the natural karyotype and the ancestral genome.

New knowledge of synteny and collinearity promises to unify genetics and to affect our perception of higher order genome structure. This exciting new synthetic approach emphasizes genomic similarities rather than diversity. Two other aspects of genomic form and organisation, offering potentially unifying concepts in genome studies are: the nucleotype, and the natural karyotype. Genome size varies greatly between eukaryotes, and shows many strikingly precise correlations with phenotypic characters, independent of information encoded in DNA. Such nucleotypic correlations, based on biophysical absolutes, apply to all species, irrespective of genome size or chromosome number, and set limits on the range of phenotypes which can be expressed by genic control. Thus, knowledge of nucleotypic effects has considerable predictive value which can help to unify our understanding of genomes. Other studies of reconstructed nuclei have shown that: (1) the basic haploid genome exists as a real structural unit in nuclear architecture; while (2) the mean spatial arrangement of its heterologues also exists as a natural karyotype which is predictable using a simple model. Recently reported conceptual alignments of the maize genomes, which reflect the circularized ancestral grass genome, show interesting similarities with the orders of centromeres in their natural karyotypes predicted by the Bennett model. The basis of this phenomenon (if repeated in other species), and of selection which retains the ancestral genome form despite changes in basic chromosome number, may need to be explained. Perhaps the overall 3-D structure of the genome has some critical functional significance, essential for development. If so, a knowledge of this common structure would further unify our understanding of genomes and their evolution.

Biological Evolution↗

A cross-comparison of a large dataset of genes.

SUMMARY: We make available a large cross-comparison for 16 of the completely sequenced genomes and additional eukaryotic genes. The alignments were performed at the protein level using liberal similarity bounds in order to capture as many significant alignments as possible. This dataset will be updated as new genomes become available.

Animals↗

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Databases, Genetic↗

Recognizing the pseudogenes in bacterial genomes.

Pseudogenes are now known to be a regular feature of bacterial genomes and are found in particularly high numbers within the genomes of recently emerged bacterial pathogens. As most pseudogenes are recognized by sequence alignments, we use newly available genomic sequences to identify the pseudogenes in 11 genomes from 4 bacterial genera, each of which contains at least 1 human pathogen. The numbers of pseudogenes range from 27 in Staphylococcus aureus MW2 to 337 in Yersinia pestis CO92 (e.g. 1-8% of the annotated genes in the genome). Most pseudogenes are formed by small frameshifting indels, but because stop codons are A + T-rich, the two low-G + C Gram-positive taxa (Streptococcus and Staphylococcus) have relatively high fractions of pseudogenes generated by nonsense mutations when compared with more G + C-rich genomes. Over half of the pseudogenes are produced from genes whose original functions were annotated as 'hypothetical' or 'unknown'; however, several broadly distributed genes involved in nucleotide processing, repair or replication have become pseudogenes in one of the sequenced Vibrio vulnificus genomes. Although many of our comparisons involved closely related strains with broadly overlapping gene inventories, each genome contains a largely unique set of pseudogenes, suggesting that pseudogenes are formed and eliminated relatively rapidly from most bacterial genomes.

Genome, Bacterial↗

Multiple sequence alignment accuracy and evolutionary distance estimation.

BACKGROUND: Sequence alignment is a common tool in bioinformatics and comparative genomics. It is generally assumed that multiple sequence alignment yields better results than pair wise sequence alignment, but this assumption has rarely been tested, and never with the control provided by simulation analysis. This study used sequence simulation to examine the gain in accuracy of adding a third sequence to a pair wise alignment, particularly concentrating on how the phylogenetic position of the additional sequence relative to the first pair changes the accuracy of the initial pair's alignment as well as their estimated evolutionary distance. RESULTS: The maximal gain in alignment accuracy was found not when the third sequence is directly intermediate between the initial two sequences, but rather when it perfectly subdivides the branch leading from the root of the tree to one of the original sequences (making it half as close to one sequence as the other). Evolutionary distance estimation in the multiple alignment framework, however, is largely unrelated to alignment accuracy and rather is dependent on the position of the third sequence; the closer the branch leading to the third sequence is to the root of the tree, the larger the estimated distance between the first two sequences. CONCLUSION: The bias in distance estimation appears to be a direct result of the standard greedy progressive algorithm used by many multiple alignment methods. These results have implications for choosing new taxa and genomes to sequence when resources are limited.

Algorithms↗

Genome-wide sequence and functional analysis of early replicating DNA in normal human fibroblasts.

BACKGROUND: The replication of mammalian genomic DNA during the S phase is a highly coordinated process that occurs in a programmed manner. Recent studies have begun to elucidate the pattern of replication timing on a genomic scale. Using a combination of experimental and computational techniques, we identified a genome-wide set of the earliest replicating sequences. This was accomplished by first creating a cosmid library containing DNA enriched in sequences that replicate early in the S phase of normal human fibroblasts. Clone ends were then sequenced and aligned to the human genome. RESULTS: By clustering adjacent or overlapping early replicating clones, we identified 1759 "islands" averaging 100 kb in length, allowing us to perform the most detailed analysis to date of DNA characteristics and genes contained within early replicating DNA. Islands are enriched in open chromatin, transcription related elements, and Alu repetitive elements, with an underrepresentation of LINE elements. In addition, we found a paucity of LTR retroposons, DNA transposon sequences, and an enrichment in all classes of tandem repeats, except for dinucleotides. CONCLUSION: An analysis of genes associated with islands revealed that nearly half of all genes in the WNT family, and a number of genes in the base excision repair pathway, including four of ten DNA glycosylases, were associated with island sequences. Also, we found an overrepresentation of members of apoptosis-associated genes in very early replicating sequences from both fibroblast and lymphoblastoid cells. These data suggest that there is a temporal pattern of replication for some functionally related genes.

Cell Proliferation↗

MAASE: an alternative splicing database designed for supporting splicing microarray applications.

Alternative splicing is a prominent feature of higher eukaryotes. Understanding of the function of mRNA isoforms and the regulation of alternative splicing is a major challenge in the post-genomic era. The development of mRNA isoform sensitive microarrays, which requires precise splice-junction sequence information, is a promising approach. Despite the availability of a large number of mRNAs and ESTs in various databases and the efforts made to align transcript sequences to genomic sequences, existing alternative splicing databases do not offer adequate information in an appropriate format to aid in splicing array design. Here we describe our effort in constructing the Manually Annotated Alternatively Spliced Events (MAASE) database system, which is specifically designed to support splicing microarray applications. MAASE comprises two components: (1) a manual/computational annotation tool for the efficient extraction of critical sequence and functional information for alternative splicing events and (2) a user-friendly database of annotated events that allows convenient export of information to aid in microarray design and data analysis. We provide a detailed introduction and a step-by-step user guide to the MAASE database system to facilitate future large-scale annotation efforts, integration with other alternative splicing databases, and splicing array fabrication.

Alternative Splicing↗

Whole-Genome Sequencing of 54 Dengchuan Cattle (Bos taurus) from Southwest China.

Domestic cattle (Bos taurus) play a significant role in human society as they provide abundant food resources and contribute to the development of agriculture and traditional culture. Dengchuan cattle, a local breed from Yunnan, Southwest China, are known for their high-quality milk and are at risk of extinction due to crossbreeding. To preserve the superior genetic resources of Dengchuan cattle, this study conducted whole-genome sequencing of 54 Dengchuan cattle using blood DNA samples, generating approximately 3.56 TB of clean data with an average sequencing depth of 32.78X. The sequencing data were aligned to the bovine reference genome (ARS-UCD2.0), achieving an average alignment rate of 99.85%. A total of 9,950,420 SNPs and 2,476,207 indels were detected using variant calling workflow. These data were utilized to characterize genomic profile of this unique cattle breed. The data generated in this study can be incorporated into the global cattle genomic diversity database, providing valuable information for comparative studies on cattle.

Animals↗

Cloning of the gene coding for a human receptor for formyl peptides. Characterization of a promoter region and evidence for polymorphic expression.

Recently we reported that, in HL-60 cells, transcription of the formyl peptide receptor (FPR) gene can be up- and downregulated by agents that induce differentiation of HL-60 cells into neutrophils. To begin studying the mechanisms involved in regulation of FPR gene expression, we cloned two human cDNAs and the gene coding for FPR. The genomic clone (pINF14) contained a 14.5-kb insert. A 2.7-kb EcoRI fragment was obtained from pINF14 that hybridized with an FPR open reading frame probe. The EcoRI fragment was sequenced and found to contain an intronless FPR open reading frame. Sequence alignment of the EcoRI genomic fragment with the FPR cDNA revealed that the first 31 bases of 5' untranslated FPR cDNA were not represented in the genomic fragment. Furthermore, a splicing consensus sequence was present in the genomic fragment at the site of divergence with the cDNA sequence. Restriction mapping and Southern blot analysis identified a 121-bp fragment that contained the sequence corresponding to the first 31 bases of 5' untranslated FPR cDNA. An additional (previously undescribed) 15-bp cDNA sequence in the 5' end of FPR were identified using an anchored polymerase chain reaction. This sequence was also contained in the genomic 121-bp fragment. This 121-bp fragment was located 5.2 kb (intron) upstream of the FPR open reading frame. It contained an unusual TATA box and displayed transcriptional activity in vitro and in vivo. Potential binding sites for AP-1 and glucocorticoid receptor were identified upstream of the putative TATA box.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

Overexpression of LRP12, a gene contained within an 8q22 amplicon identified by high-resolution array CGH analysis of oral squamous cell carcinomas.

Chromosome 8q amplification is a common event observed in cancer. In this study, we used high-resolution array comparative genomic hybridization to resolve two neighboring regions on 8q that are both amplified in oral cancer. One region (at 8q24) contains the MYC oncogene, which is frequently overexpressed in many cancers, while the other region (at 8q22) represents a novel amplicon. The alignment of array comparative genomic hybridization profiles of 20 microdissected oral squamous cell carcinomas (OSCCs) revealed a approximately 5 Mbp region of frequent copy number alteration. This region harbors 16 known genes. Gene expression analysis comparing 15 microdissected OSCC with 16 normal epithelium samples revealed overexpression specific to LRP12 but not the neighboring genes, dihydropyrimidinase and FOG2, suggesting that LRP12 may function as an oncogene in oral tumors.

Biomarkers, Tumor↗