Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Leveraging the mouse genome for gene prediction in human: from whole-genome shotgun reads to a global synteny map.

The availability of draft sequences for both the mouse and human genomes makes it possible, for the first time, to annotate whole mammalian genomes using comparative methods. TWINSCAN is a gene-prediction system that combines the methods of single-genome predictors like GENSCAN with information derived from genome comparison, thereby improving accuracy. Because TWINSCAN uses genomic sequence only, it is less biased toward highly and/or ubiquitously expressed genes than GENEWISE, GENOMESCAN, and other methods based on evidence derived from transcripts. We show that TWINSCAN improves gene prediction in human using intermediate products from various stages of the sequencing and analysis of the mouse genome, from low-redundancy, whole-genome shotgun reads to the draft assembly and the synteny map. TWINSCAN improves on the prior state of the art even when alignments from only 1X coverage of the mouse genome are available. Gene prediction accuracy improves steadily from 1X through 3X, more slowly from 3X to 4X, and relatively little thereafter. The assembly and the synteny map greatly speed the computations, however. Our human annotation using the mouse assembly is conservative, predicting only 25,622 genes, and appears to be one of the best de novo annotations of the human genome to date.

Animals↗

Genome-Wide Identification and Bioinformatics Analysis of the FAD Gene Family in Walnut (Juglans regia L.).

Fatty acid desaturase (FAD) is a core catalytic enzyme in plants for the synthesis of unsaturated fatty acids, profoundly affecting plant growth, development, and adaptability to various environmental stresses. The walnut (Juglans regia L.) is an important woody oil tree species, and its kernel is rich in unsaturated fatty acids. Systematic identification of the walnut FAD gene family and analysis of its function are of great significance for revealing the molecular mechanisms underlying unsaturated fatty acid metabolism in the walnut. Based on walnut whole-genome data, this study used homology alignment and hidden Markov model search methods to identify the JrFAD gene family members. Subsequently, a variety of bioinformatics tools were used to systematically analyze their structural characteristics, evolutionary expansion mechanism, expression regulation, and function. A total of 21 JrFAD gene family members were identified and classified into five subfamilies. The family genes were unevenly distributed on nine chromosomes. WGD/segmental duplication was the main expansion method, and the duplicated gene pairs experienced strong purification selection. The family gene promoter sequence is rich in regulatory elements that respond to light, plant hormones, and various stresses. The expression pattern analysis showed that JrFAD3.1 and JrFAD2.3 showed high expression specifically during the rapid accumulation of walnut kernel oil. This study clarified the composition and evolutionary characteristics of the FAD gene family in the walnut, which provides useful information for in-depth analyses of its functional mechanism in the regulation of lipid metabolism, and also identified potential candidate gene resources for the genetic improvement of walnut varieties with high amounts of unsaturated fatty acids.

Juglans↗

EVOPRINTER, a multigenomic comparative tool for rapid identification of functionally important DNA.

Here, we describe a multigenomic DNA sequence-analysis tool, evoprinter, that facilitates the rapid identification of evolutionary conserved sequences within the context of a single species. The evoprinter output identifies multispecies-conserved DNA sequences as they exist in a reference DNA. This identification is accomplished by superimposing multiple reference DNA vs. test-genome pairwise blat (blast-like alignment tool) readouts of the reference DNA to identify conserved nucleotides that are shared by all orthologous DNAs. evoprinter analysis of well characterized genes reveals that most, if not all, of the conserved sequences are essential for gene function. For example, analysis of orthologous genes that are shared by many vertebrates identifies conserved DNA in both protein-encoding sequences and noncoding cis-regulatory regions, including enhancers and mRNA microRNA binding sites. In Drosophila, the combined mutational histories of five or more species affords near-base pair resolution of conserved transcription factor DNA-binding sites, and essential amino acids are revealed by the nucleotide flexibility of their codon-wobble position(s). Conserved small peptide-encoding genes, which had been undetected by conventional gene-prediction algorithms, are identified by the codon-wobble signatures of invariant amino acids. Also, evoprinter allows one to assess the degree of evolutionary divergence between orthologous DNAs by highlighting differences between a selected species and the other test species.

Animals↗

Clustering protein sequences with a novel metric transformed from sequence similarity scores and sequence alignments with neural networks.

BACKGROUND: The sequencing of the human genome has enabled us to access a comprehensive list of genes (both experimental and predicted) for further analysis. While a majority of the approximately 30,000 known and predicted human coding genes are characterized and have been assigned at least one function, there remains a fair number of genes (about 12,000) for which no annotation has been made. The recent sequencing of other genomes has provided us with a huge amount of auxiliary sequence data which could help in the characterization of the human genes. Clustering these sequences into families is one of the first steps to perform comparative studies across several genomes. RESULTS: Here we report a novel clustering algorithm (CLUGEN) that has been used to cluster sequences of experimentally verified and predicted proteins from all sequenced genomes using a novel distance metric which is a neural network score between a pair of protein sequences. This distance metric is based on the pairwise sequence similarity score and the similarity between their domain structures. The distance metric is the probability that a pair of protein sequences are of the same Interpro family/domain, which facilitates the modelling of transitive homology closure to detect remote homologues. The hierarchical average clustering method is applied with the new distance metric. CONCLUSION: Benchmarking studies of our algorithm versus those reported in the literature shows that our algorithm provides clustering results with lower false positive and false negative rates. The clustering algorithm is applied to cluster several eukaryotic genomes and several dozens of prokaryotic genomes.

Algorithms↗

Comparison of the di- and trinucleotide frequencies from the genomes of nine different coronaviruses.

As an alternative to protein alignments for the comparison of sequences, the reiterations of mono- di- and trinucleotide frequencies were used for the comparison of coronavirus sequences. The relative abundance of the di- and trinucleotide frequencies within the 3' part from nine coronavirus genomes were determined. The patterns of dinucleotide frequencies and the trinucleotide frequencies showed some common features for all coronaviruses but also differences between the groups formerly defined on the base of antigenic relatedness. The normalised dinucleotide frequencies were further used to calculate the distances between coronavirus sequences. Based on the dinucleotide frequency distances, coronaviruses can be divided into two groups which roughly reflect the taxonomic groups. In this kind of evaluation, however, IBV occupies a position different to the one that it would take based on most protein sequence comparisons. Based on similarities within coding sequences and antigenic properties IBV occupies a place outside of both groups. Based on the dinucleotide frequencies IBV gained a position in between of the TGEV-related and the MHV-clustered coronaviruses.

Animals↗

Using ESTs to improve the accuracy of de novo gene prediction.

BACKGROUND: ESTs are a tremendous resource for determining the exon-intron structures of genes, but even extensive EST sequencing tends to leave many exons and genes untouched. Gene prediction systems based exclusively on EST alignments miss these exons and genes, leading to poor sensitivity. De novo gene prediction systems, which ignore ESTs in favor of genomic sequence, can predict such "untouched" exons, but they are less accurate when predicting exons to which ESTs align. TWINSCAN is the most accurate de novo gene finder available for nematodes and N-SCAN is the most accurate for mammals, as measured by exact CDS gene prediction and exact exon prediction. RESULTS: TWINSCAN_EST is a new system that successfully combines EST alignments with TWINSCAN. On the whole C. elegans genome TWINSCAN_EST shows 14% improvement in sensitivity and 13% in specificity in predicting exact gene structures compared to TWINSCAN without EST alignments. Not only are the structures revealed by EST alignments predicted correctly, but these also constrain the predictions without alignments, improving their accuracy. For the human genome, we used the same approach with N-SCAN, creating N-SCAN_EST. On the whole genome, N-SCAN_EST produced a 6% improvement in sensitivity and 1% in specificity of exact gene structure predictions compared to N-SCAN. CONCLUSION: TWINSCAN_EST and N-SCAN_EST are more accurate than TWINSCAN and N-SCAN, while retaining their ability to discover novel genes to which no ESTs align. Thus, we recommend using the EST versions of these programs to annotate any genome for which EST information is available.TWINSCAN_EST and N-SCAN_EST are part of the TWINSCAN open source software package http://genes.cse.wustl.edu/distribution/download_TS.html.

Algorithms↗

GenoMycDB: a database for comparative analysis of mycobacterial genes and genomes.

Several databases and computational tools have been created with the aim of organizing, integrating and analyzing the wealth of information generated by large-scale sequencing projects of mycobacterial genomes and those of other organisms. However, with very few exceptions, these databases and tools do not allow for massive and/or dynamic comparison of these data. GenoMycDB (http://www.dbbm.fiocruz.br/GenoMycDB) is a relational database built for large-scale comparative analyses of completely sequenced mycobacterial genomes, based on their predicted protein content. Its central structure is composed of the results obtained after pair-wise sequence alignments among all the predicted proteins coded by the genomes of six mycobacteria: Mycobacterium tuberculosis (strains H37Rv and CDC1551), M. bovis AF2122/97, M. avium subsp. paratuberculosis K10, M. leprae TN, and M. smegmatis MC2 155. The database stores the computed similarity parameters of every aligned pair, providing for each protein sequence the predicted subcellular localization, the assigned cluster of orthologous groups, the features of the corresponding gene, and links to several important databases. Tables containing pairs or groups of potential homologs between selected species/strains can be produced dynamically by user-defined criteria, based on one or multiple sequence similarity parameters. In addition, searches can be restricted according to the predicted subcellular localization of the protein, the DNA strand of the corresponding gene and/or the description of the protein. Massive data search and/or retrieval are available, and different ways of exporting the result are offered. GenoMycDB provides an on-line resource for the functional classification of mycobacterial proteins as well as for the analysis of genome structure, organization, and evolution.

Bacterial Proteins↗

SPEM: improving multiple sequence alignment with sequence profiles and predicted secondary structures.

MOTIVATION: Multiple sequence alignment is an essential part of bioinformatics tools for a genome-scale study of genes and their evolution relations. However, making an accurate alignment between remote homologs is challenging. Here, we develop a method, called SPEM, that aligns multiple sequences using pre-processed sequence profiles and predicted secondary structures for pairwise alignment, consistency-based scoring for refinement of the pairwise alignment and a progressive algorithm for final multiple alignment. RESULTS: The alignment accuracy of SPEM is compared with those of established methods such as ClustalW, T-Coffee, MUSCLE, ProbCons and PRALINE(PSI) in easy (homologs) and hard (remote homologs) benchmarks. Results indicate that the average sum of pairwise alignment scores given by SPEM are 7-15% higher than those of the methods compared in aligning remote homologs (sequence identity <30%). Its accuracy for aligning homologs (sequence identity >30%) is statistically indistinguishable from those of the state-of-the-art techniques such as ProbCons or MUSCLE 6.0. AVAILABILITY: The SPEM server and its executables are available on http://theory.med.buffalo.edu.

Algorithms↗

Hydropathy profile alignment: a tool to search for structural homologues of membrane proteins.

Hydropathy profile alignment is introduced as a tool in functional genomics. The architecture of membrane proteins is reflected in the hydropathy profile of the amino acid sequence. Both secondary and tertiary structural elements determine the profile which provides enough sensitivity to detect evolutionary links between membrane proteins that are based on structural rather than sequence similarities. Since structure is better conserved than amino acid sequence, the hydropathy profile can detect more distant evolutionary relationships than can be detected by the primary structure. The technique is demonstrated by two approaches in the analysis of a subset of membrane proteins coded on the Escherichia coli and Bacillus subtilis genomes. The subset includes secondary transporters of the 12 helix type. In the first approach, the hydropathy profiles of proteins for which no function is known are aligned with the profiles of all other proteins in the subset to search for structural paralogues with known function. In the second approach, family hydropathy profiles of 8 defined families of secondary transporters that fall into 4 different structural classes (SC-ST1-4) are used to screen the membrane protein set for members of the structural classes. The analysis reveals that over 100 membrane proteins on each genome fall in only two structural classes. The largest structural class, SC-ST1, correlates largely with the Major Facilitator Superfamily defined before, but the number of families within the class has increased up to 57. The second large structural class, SC-ST2 contains secondary transporters for amino acids and amines and consists of 12 families.

Amino Acid Sequence↗

Construction and analysis of a human-chimpanzee comparative clone map.

The recently released human genome sequences provide us with reference data to conduct comparative genomic research on primates, which will be important to understand what genetic information makes us human. Here we present a first-generation human-chimpanzee comparative genome map and its initial analysis. The map was constructed through paired alignment of 77,461 chimpanzee bacterial artificial chromosome end sequences with publicly available human genome sequences. We detected candidate positions, including two clusters on human chromosome 21 that suggest large, nonrandom regions of difference between the two genomes.

Animals↗

Identification and molecular analysis of a third Aspergillus nidulans alcohol dehydrogenase gene.

An Aspergillus nidulans functional cDNA encoding an alcohol dehydrogenase (ADH) was isolated by its ability to complement an adh1 mutation in Saccharomyces cerevisiae. Alignment of the cDNA and cloned genomic DNA sequences indicated that the ADH gene contains two small introns. The presence of ethanol in the growth medium was shown to result in ADH mRNA accumulation presumably due to transcriptional induction of the gene. However, ADH mRNA accumulation was at most only partially repressed by the presence of glucose. The ADH gene characterized here is designated ADH3 since it is distinct from the alcA gene which encodes ADH I and appears distinct from the gene which encodes ADH II. We demonstrated that the first intron in the A. nidulans ADH3 gene was not efficiently spliced in S. cerevisiae whereas the promoter region was utilized weakly. We also present a comparison of the primary structure of A. nidulans ADH III with the alcohol dehydrogenases of S. cerevisiae and Schizosaccharomyces pombe.

Alcohol Dehydrogenase↗

Toxoplasma gondii: structure and characterization of the 26S ribosomal RNA and peptidyl transferase domain.

The 26S, 5.8S, and the intergenic spacer ribosomal DNAs (rDNA) of Toxoplasma gondii have been cloned and completely sequenced from both DNA strands. The length of the large subunit was found to be 3487 nucleotides and the 5.8S was 153 nucleotides long. These formed the large rRNA subunit of Toxoplasma and were mapped in the rDNA unit known to be repeated 110 times in a head-to-tail fashion in the genome. Primer extension analysis and multiple alignments localized the 5' end point of the two rRNAs. Comparisons with Toxoplasma rDNA by nucleic acid homology studies gave 76% similarity with the dinoflagellate Prorocentrum micans, 66% with yeast Saccharomyces cerevisiae, and 64% with the ciliate Tetrahymena thermophila. Similarity was apparent in the conserved core structure of the large subunit rRNA and divergent sequences were identified in the so-called divergent domains. Construction of a secondary structure model of the peptidyl transferase center of the large rRNA revealed similarities with the same domain from other life forms.

Animals↗

Genome resources and comparative analysis tools for cardiovascular research.

Disorders of the cardiovascular system are often caused by the interaction of genetic and environmental factors that jointly contribute to individual susceptibility. Genomic data and bioinformatics tools generated from genome projects, coupled with functional verification, offer novel approaches to study both rare single-gene and complex multigenic cardiovascular diseases. These approaches include gene mapping using genome variation, especially single-nucleotide polymorphisms and comparative genomics within and between species. This chapter illustrates the major genome resources, associated bioinformatics tools, and their potential application in cardiovascular research.

Cardiovascular Diseases↗

Evolution of the recA gene and the molecular phylogeny of bacteria.

The DNA sequences of the recA gene from 25 strains of bacteria are known. The evolution of these recA gene sequences, and of the derived RecA protein sequences, is examined, with special reference to the effect of variations in genomic G + C content. From the aligned RecA protein sequences, phylogenetic trees have been drawn using both distance matrix and maximum parsimony methods. There is a broad concordance between these trees and those derived from other data (largely 16S ribosomal RNA sequences). There is a fair degree of certainty in the relationships among the "Purple" or Proteobacteria, but the branching pattern between higher taxa within the eubacteria cannot be reliably resolved with these data.

Amino Acid Sequence↗

Quantitative trait loci for lodging resistance, plant height and partial resistance to mycosphaerella blight in field pea (Pisum sativum L.).

With the development of genetic maps and the identification of the most-likely positions of quantitative trait loci (QTLs) on these maps, molecular markers for lodging resistance can be identified. Consequently, marker-assisted selection (MAS) has the potential to improve the efficiency of selection for lodging resistance in a breeding program. This study was conducted to identify genetic loci associated with lodging resistance, plant height and reaction to mycosphaerella blight in pea. A population consisting of 88 recombinant inbred lines (RILs) was developed from a cross between Carneval and MP1401. The RILs were evaluated in 11 environments across the provinces of Manitoba, Saskatchewan and Alberta, Canada in 1998, 1999 and 2000. One hundred and ninety two amplified fragment length polymorphism (AFLP) markers, 13 random amplified polymorphic DNA (RAPD) markers and one sequence tagged site (STS) marker were assigned to ten linkage groups (LGs) that covered 1,274 centi Morgans (cM) of the pea genome. Six of these LGs were aligned with the previous pea map. Two QTLs were identified for lodging resistance that collectively explained 58% of the total phenotypic variation in the mean environment. Three QTLs were identified each for plant height and resistance to mycosphaerella blight, which accounted for 65% and 36% of the total phenotypic variation, respectively, in the mean environment. These QTLs were relatively consistent across environments. The AFLP marker that was associated with the major locus for lodging resistance was converted into the sequence-characterized amplified-region (SCAR) marker. The presence or absence of the SCAR marker corresponded well with the lodging reaction of 50 commercial pea varieties.

Ascomycota↗

The role of context-dependent mutations in generating compositional and codon usage bias in grass chloroplast DNA.

The influence of local base composition on mutations in chloroplast DNA (cpDNA) is studied in detail and the resulting, empirically derived, mutation dynamics are used to analyze both base composition and codon usage bias. A 4 x 4 substitution matrix is generated for each of the 16 possible flanking base combinations (contexts) using 17,253 noncoding sites, 1309 of which are variable, from an alignment of three complete grass chloroplast genome sequences. It is shown that substitution bias at these sites is correlated with flanking base composition and that the A+T content of these flanking sites as well as the number of flanking pyrimidines on the same strand appears to have general influences on substitution properties. The context-dependent equilibrium base frequencies predicted from these matrices are then applied to two analyses. The first examines whether or not context dependency of mutations is sufficient to generate average compositional differences between noncoding cpDNA and silent sites of coding sequences. It is found that these two classes of sites exist, on average, in very different contexts and that the observed mutation dynamics are expected to generate significant differences in overall composition bias that are similar to the differences observed in cpDNA. Context dependency, however, cannot account for all of the observed differences: although silent sites in coding regions appear to be at the equilibrium predicted, noncoding cpDNA has a significantly lower A+T content than expected from its own substitution dynamics, possibly due to the influence of indels. The second study examines the codon usage of low-expression chloroplast genes. When context is accounted for, codon usage is very similar to what is predicted by the substitution dynamics of noncoding cpDNA. However, certain codon groups show significant deviation when followed by a purine in a manner suggesting some form of weak selection other than translation efficiency. Overall, the findings indicate that a full understanding of mutational dynamics is critical to understanding the role selection plays in generating composition bias and sequence structure.

Base Composition↗

Evaluation of five ab initio gene prediction programs for the discovery of maize genes.

Five ab initio programs (FGENESH, GeneMark.hmm, GENSCAN, GlimmerR and Grail) were evaluated for their accuracy in predicting maize genes. Two of these programs, GeneMark.hmm and GENSCAN had been trained for maize; FGENESH had been trained for monocots (including maize), and the others had been trained for rice or Arabidopsis. Initial evaluations were conducted using eight maize genes (gl8a, pdc2, pdc3, rf2c, rf2d, rf2e1, rth1, and rth3) of which the sequences were not released to the public prior to conducting this evaluation. The significant advantage of this data set for this evaluation is that these genes could not have been included in the training sets of the prediction programs. FGENESH yielded the most accurate and GeneMark.hmm the second most accurate predictions. The five programs were used in conjunction with RT-PCR to identify and establish the structures of two new genes in the a1-sh2 interval of the maize genome. FGENESH, GeneMark.hmm and GENSCAN were tested on a larger data set consisting of maize assembled genomic islands (MAGIs) that had been aligned to ESTs. FGENESH, GeneMark.hmm and GENSCAN correctly predicted gene models in 773, 625, and 371 MAGIs, respectively, out of the 1353 MAGIs that comprise data set 2.

Alternative Splicing↗

Effects of CRISPR technology on agricultural sustainability: global applications and turkish perspective.

This review evaluates CRISPR/Cas applications in agriculture from a global perspective with explicit reference to T&#xfc;rkiye. Using a literature gap-matrix approach organised around four analytical dimensions-environmental, economic, social and policy, and scientific and technological-we synthesize the primary evidence on water and input use, productivity, disease resistance, and product quality. The literature concentrates on water and fertilizer use, productivity, and off-target accuracy, whereas soil health, biodiversity, consumer acceptance, ethical considerations and regulatory frameworks remain systematically under-represented. Global deployment of CRISPR is already delivering measurable advantages in food security, shelf life and nutritional value, while in T&#xfc;rkiye the research base is at an early stage but has clear potential in wheat, barley, tomato and olive. Translating CRISPR into Turkish agricultural sustainability requires (i) a domestic biosafety framework aligned with the emerging European New Genomic Techniques approach, (ii) sustained investment in multi-location primary field trials, and (iii) inclusive deployment mechanisms-particularly through producer cooperatives-that allow smallholder farmers to benefit from edited varieties.

Plants, Genetically Modified↗