Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

VaZyMolO: a tool to define and classify modularity in viral proteins.

Viral structural genomic projects aim at unveiling the function of unknown viral proteins by employing high-throughput approaches to determine their 3D structure and to identify their function through fold-homology studies. The 'viral enzyme module localization' (VaZyMolO) tool has been developed, which aims at defining viral protein modules that might be expressed in a soluble and functionally active form, thereby identifying candidates for crystallization studies. VaZyMolO includes 114 complete viral genome sequences of both negative- and positive-sense, single-stranded RNA viruses available from NCBI. In VaZyMolO, a module is defined as a structural and/or functional unit. Modules were first identified by homology search and then validated by the convergence of results from sequence composition analysis, motif search, transmembrane region search and domain definitions, as found in the literature. The public interface of VaZyMolO, which is accessible from http://www.vazymolo.org, allows comparison of a query sequence to all VaZyMolO modules of known function.

DNA, Complementary↗

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article↗

Links from genome proteins to known 3-D structures.

We describe a genome annotation service provided by the Entrez browser, http://www.ncbi.nlm.nih.gov/entrez. All protein products identified in fully sequenced microbial genomes have been compared with proteins with known 3-D structure by use of the BLAST sequence comparison algorithm. For the approximately 20% of genome proteins in which unambiguous sequence similarity is detected, Entrez provides a link from the gene product to its predicted structure. The service uses the Cn3D molecular graphics viewer to present a 3-D view of the known structure, together with an alignment display mapping conserved residues from the genome protein onto the known structure. Using an example from Aeropyrum pernix, we illustrate how mapping to a 3-D structure can confirm predictions of biological function.

Algorithms↗

Epidermodysplasia verruciformis in a patient with Hodgkin's disease: characterization of a new papillomavirus type and interferon treatment.

A new human papillomavirus (HPV) was discovered in disseminated, macular, pityriasis versicolor-like lesions on the skin of the neck, face, scalp, and pubic region of a 42-year-old male suffering from Hodgkin's disease. Histopathology revealed features characteristic of epidermodysplasia verruciformis (ev). In contrast to classical ev, the lesions were almost exclusively seen in previously irradiated and UV-exposed skin areas. Papillomavirus capsid antigen was demonstrated with the genus-specific antiserum and the patient's serum, which had IgM and IgG antibody titers. HPV DNA was isolated from biopsies and cloned into the vector pIC20H. It proved to be related to ev-associated viruses, showing 23% cross-hybridization with DNA of the closest relative HPV14. The new HPV type was named HPV46. The genome was physically mapped and colinearly aligned with HPV8 DNA to establish its gene organization. Interferon treatment of the patient did not significantly change the clinical picture nor was the concentration of viral DNA per lesion affected. However, no virus capsid antigen was detectable after starting treatment.

Adult↗

Stress-dependent expression of a polymorphic, charged antigen in the protozoan parasite Entamoeba histolytica.

We have identified a novel stress inducible gene, Ehssp1 in Entamoeba histolytica, the causative agent of amebiasis. Ehssp1 belongs to a polymorphic, multigene family and is present on multiple chromosomes. No homologue of this gene was found in the NCBI database. Sequence alignment of the multiple copies, and genomic PCR data restricted the polymorphism to the central region of the gene. This region contains a polypurine stretch that encodes a domain rich in acidic and basic amino acids. Under normal culture conditions only one copy of this multigene family is expressed, as observed by Northern blot and RT-PCR analysis. The size of this copy of the gene is 1,077 nucleotides, encoding a protein of 359 amino acids. The polymorphic domain in this copy is 64 nucleotides long. However, on exposure of cells to stress conditions such as heat shock or oxidative stress, multiple polymorphic copies of the gene are expressed, suggesting a possible role of this gene in adaptation of cells to stress conditions. The gene copy expressed under normal conditions, and the expression profile of cells under heat stress was identical in two different strains of E. histolytica tested. Interestingly, the extent of polymorphism in this gene was very less in E. dispar, a nonpathogenic sibling species of E. histolytica. Ehssp1 was found to be antigenic in invasive amebiasis patients.

Amino Acid Sequence↗

nirA, the pathway-specific regulatory gene of nitrate assimilation in Aspergillus nidulans, encodes a putative GAL4-type zinc finger protein and contains four introns in highly conserved regions.

The nucleotide sequence of nirA, mediating nitrate induction in Aspergillus nidulans, has been determined. Alignment of the cDNA and the genomic DNA sequence indicates that the gene contains four introns and encodes a protein of 892 amino acids. The deduced NIRA protein displays all characteristics of a transcriptional activator. A putative double-stranded DNA-binding domain in the amino-terminal part comprises six cysteine residues, characteristic for the GAL4 family of zinc finger proteins. An amino-terminal highly acidic region and two proline-rich regions are also present. The nucleotide sequences of two mutations were determined after they were mapped by transformation with overlapping DNA fragments, amplified by the polymerase chain reaction. nirA87, a mutation conferring noninducibility by nitrate and nitrite, has a -1 frameshift at triplet 340, which eliminates 549 C-terminal amino acids from the polypeptide. Under the assumption that the truncated polypeptide is stable, it comprises the zinc finger domain and the acidic region, which seem not sufficient for transcriptional activation. nirAd-106, an allele conferring nitrogen metabolite derepression of nitrate and nitrite reductase activity, includes two transitions, changing a glutamic acid to a lysine and a valine to an alanine, situated between a basic and a proline-rich region of the protein. Northern (RNA) analysis of the wild type and of constitutive (nirAc) and derepressed (nirAd) mutants show that the nirA transcript does not vary between these strains, being in all cases constitutively expressed. On the other hand, transcript levels of structural genes (niaD and niiA) do vary, being highly inducible in the wild type but constitutively expressed in the nirAc mutant. The nirAd mutant appears phenotypically derepressed, because the niaD and niiA transcript levels are overinduced in the presence of nitrate but are still partially repressed in the presence of ammonium.

Amino Acid Sequence↗

Systematic screening of sheep skin cDNA libraries for microsatellite sequences.

65,000 sheep skin cDNA clones were gridded in high density on to nylon membranes and screened for (CA)n and (GA)n repeat containing clones. 296 dinucleotide repeat-containing clones were identified with approximately 85% non-redundancy. Clones were single-pass 5' sequenced and we compared the Expressed Sequence Tag (EST) sequences to the Swiss-Prot database to ascertain their identity and/or putative function. We then aligned the ESTs against the human genomic sequence to determine the locations of human orthologous sequences. Finally, we developed a subset of polymorphic microsatellite markers and positioned them on the ovine linkage map.

Animals↗

NemaFootPrinter: a web based software for the identification of conserved non-coding genome sequence regions between C. elegans and C. briggsae.

BACKGROUND: NemaFootPrinter (Nematode Transcription Factor Scan Through Philogenetic Footprinting) is a web-based software for interactive identification of conserved, non-exonic DNA segments in the genomes of C. elegans and C. briggsae. It has been implemented according to the following project specifications:a) Automated identification of orthologous gene pairs. b) Interactive selection of the boundaries of the genes to be compared. c) Pairwise sequence comparison with a range of different methods. d) Identification of putative transcription factor binding sites on conserved, non-exonic DNA segments. RESULTS: Starting from a C. elegans or C. briggsae gene name or identifier, the software identifies the putative ortholog (if any), based on information derived from public nematode genome annotation databases. The investigator can then retrieve the genome DNA sequences of the two orthologous genes; visualize graphically the genes' intron/exon structure and the surrounding DNA regions; select, through an interactive graphical user interface, subsequences of the two gene regions. Using a bioinformatics toolbox (Blast2seq, Dotmatcher, Ssearch and connection to the rVista database) the investigator is able at the end of the procedure to identify and analyze significant sequences similarities, detecting the presence of transcription factor binding sites corresponding to the conserved segments. The software automatically masks exons. DISCUSSION: This software is intended as a practical and intuitive tool for the researchers interested in the identification of non-exonic conserved sequence segments between C. elegans and C. briggsae. These sequences may contain regulatory transcriptional elements since they are conserved between two related, but rapidly evolving genomes. This software also highlights the power of genome annotation databases when they are conceived as an open resource and the possibilities offered by seamless integration of different web services via the http protocol. AVAILABILITY: The program is freely available at http://bio.ifom-firc.it/NTFootPrinter.

Animals↗

preAssemble: a tool for automatic sequencer trace data processing.

BACKGROUND: Trace or chromatogram files (raw data) are produced by automatic nucleic acid sequencing equipment or sequencers. Each file contains information which can be interpreted by specialised software to reveal the sequence (base calling). This is done by the sequencer proprietary software or publicly available programs. Depending on the size of a sequencing project the number of trace files can vary from just a few to thousands of files. Sequencing quality assessment on various criteria is important at the stage preceding clustering and contig assembly. Two major publicly available packages--Phred and Staden are used by preAssemble to perform sequence quality processing. RESULTS: The preAssemble pre-assembly sequence processing pipeline has been developed for small to large scale automatic processing of DNA sequencer chromatogram (trace) data. The Staden Package Pregap4 module and base-calling program Phred are utilized in the pipeline, which produces detailed and self-explanatory output that can be displayed with a web browser. preAssemble can be used successfully with very little previous experience, however options for parameter tuning are provided for advanced users. preAssemble runs under UNIX and LINUX operating systems. It is available for downloading and will run as stand-alone software. It can also be accessed on the Norwegian Salmon Genome Project web site where preAssemble jobs can be run on the project server. CONCLUSION: preAssemble is a tool allowing to perform quality assessment of sequences generated by automatic sequencing equipment. preAssemble is flexible since both interactive jobs on the preAssemble server and the stand alone downloadable version are available. Virtually no previous experience is necessary to run a default preAssemble job, on the other hand options for parameter tuning are provided. Consequently preAssemble can be used as efficiently for just several trace files as for large scale sequence processing.

Algorithms↗

Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.

BACKGROUND: In order to improve gene prediction, extrinsic evidence on the gene structure can be collected from various sources of information such as genome-genome comparisons and EST and protein alignments. However, such evidence is often incomplete and usually uncertain. The extrinsic evidence is usually not sufficient to recover the complete gene structure of all genes completely and the available evidence is often unreliable. Therefore extrinsic evidence is most valuable when it is balanced with sequence-intrinsic evidence. RESULTS: We present a fairly general method for integration of external information. Our method is based on the evaluation of hints to potentially protein-coding regions by means of a Generalized Hidden Markov Model (GHMM) that takes both intrinsic and extrinsic information into account. We used this method to extend the ab initio gene prediction program AUGUSTUS to a versatile tool that we call AUGUSTUS+. In this study, we focus on hints derived from matches to an EST or protein database, but our approach can be used to include arbitrary user-defined hints. Our method is only moderately effected by the length of a database match. Further, it exploits the information that can be derived from the absence of such matches. As a special case, AUGUSTUS+ can predict genes under user-defined constraints, e.g. if the positions of certain exons are known. With hints from EST and protein databases, our new approach was able to predict 89% of the exons in human chromosome 22 correctly. CONCLUSION: Sensitive probabilistic modeling of extrinsic evidence such as sequence database matches can increase gene prediction accuracy. When a match of a sequence interval to an EST or protein sequence is used it should be treated as compound information rather than as information about individual positions.

Algorithms↗

Phylogenetic analysis of bacterial and archaeal arsC gene sequences suggests an ancient, common origin for arsenate reductase.

BACKGROUND: The ars gene system provides arsenic resistance for a variety of microorganisms and can be chromosomal or plasmid-borne. The arsC gene, which codes for an arsenate reductase is essential for arsenate resistance and transforms arsenate into arsenite, which is extruded from the cell. A survey of GenBank shows that arsC appears to be phylogenetically widespread both in organisms with known arsenic resistance and those organisms that have been sequenced as part of whole genome projects. RESULTS: Phylogenetic analysis of aligned arsC sequences shows broad similarities to the established 16S rRNA phylogeny, with separation of bacterial, archaeal, and subsequently eukaryotic arsC genes. However, inconsistencies between arsC and 16S rRNA are apparent for some taxa. Cyanobacteria and some of the gamma-Proteobacteria appear to possess arsC genes that are similar to those of Low GC Gram-positive Bacteria, and other isolated taxa possess arsC genes that would not be expected based on known evolutionary relationships. There is no clear separation of plasmid-borne and chromosomal arsC genes, although a number of the Enterobacteriales (gamma-Proteobacteria) possess similar plasmid-encoded arsC sequences. CONCLUSION: The overall phylogeny of the arsenate reductases suggests a single, early origin of the arsC gene and subsequent sequence divergence to give the distinct arsC classes that exist today. Discrepancies between 16S rRNA and arsC phylogenies support the role of horizontal gene transfer (HGT) in the evolution of arsenate reductases, with a number of instances of HGT early in bacterial arsC evolution. Plasmid-borne arsC genes are not monophyletic suggesting multiple cases of chromosomal-plasmid exchange and subsequent HGT. Overall, arsC phylogeny is complex and is likely the result of a number of evolutionary mechanisms.

Adenosine Triphosphatases↗

Can sequence determine function?

The functional annotation of proteins identified in genome sequencing projects is based on similarities to homologs in the databases. As a result of the possible strategies for divergent evolution, homologous enzymes frequently do not catalyze the same reaction, and we conclude that assignment of function from sequence information alone should be viewed with some skepticism.

Animals↗

Thermostable esterase from a thermoacidophilic archaeon: purification and characterization for enzymatic resolution of a chiral compound.

Homolog to lipolytic enzymes having the consensus sequence Gly-X-Ser-X-Gly, from the Sulfolobus solfataricus P2 genome, were identified by multiple sequence alignments. Among three potential candidate sequences, one (Est3), which displayed higher activity than the other enzymes on the indicate plates, was characterized. The gene (est 3) was expressed in Escherichia coli, and the recombinant protein (Est3) was purified by chromatographic separation. The enzyme is a trimeric protein and has a molecular weight of 32 kDa in monomer form in its native structure. The optimal pH and temperature of the esterase were 7.4 and 80 degrees C respectively. The enzyme showed broad substrate specificities toward various p-nitrophenyl esters ranging from C2 to C16. The catalytic activity of the Est3 esterase was strongly inhibited by phenylmethylsulfonyl fluoride (PMSF) and diethyl p-nitrophenyl phosphate. Based on substrate specificity and the action of inhibitors, the Est3 enzyme was estimated to be a carboxylesterase (EC 3.1.1.1). The enzyme with methyl (+/-)-2-(3-benzoylphenyl)propionate-hydrolyzing activity to (-)-2-(3-benzoylphenyl)propionic acid displayed a moderate degree of enantioselectivity. The product, (-)-2-(3-benzoylphenyl)propionic acid, rather than its methyl ester, was obtained in 80% enantiomeric excess (e.e.(p)) at 20% conversion at 60 degrees C after a 32-h reaction. This result indicates that S. solfataricus esterase can be used for application in the synthesis of chiral compounds.

Amino Acid Sequence↗

Genotypic characteristics of bovine viral diarrhea virus 2 strains isolated in northern Italy.

Two strains of Bovine viral diarrhea virus 2 (BVDV-2) were isolated from calves in northern Italy. Variations in the 5'-untranslated region (UTR) of the genome were studied by primary structure alignment and neighbor-joining method based phylogenetic tree analyses and by palindromic nucleotide substitutions at the three variable loci in the 5'-UTR. Genetic analysis indicated their appurtenance to genovar BVDV-2a. Nucleotide sequence at the 5'-UTR of strain BS-95-II, one of the Italian isolates from healthy calves, showed 98% homology to that of the Japanese isolate OY89, a cytopathic strain derived from cattle with mucosal disease.

5' Untranslated Regions↗

The uncoupling protein 1 gene (UCP1) is disrupted in the pig lineage: a genetic explanation for poor thermoregulation in piglets.

Piglets appear to lack brown adipose tissue, a specific type of fat that is essential for nonshivering thermogenesis in mammals, and they rely on shivering as the main mechanism for thermoregulation. Here we provide a genetic explanation for the poor thermoregulation in pigs as we demonstrate that the gene for uncoupling protein 1 (UCP1) was disrupted in the pig lineage. UCP1 is exclusively expressed in brown adipose tissue and plays a crucial role for thermogenesis by uncoupling oxidative phosphorylation. We used long-range PCR and genome walking to determine the complete genome sequence of pig UCP1. An alignment with human UCP1 revealed that exons 3 to 5 were eliminated by a deletion in the pig sequence. The presence of this deletion was confirmed in all tested domestic pigs, as well as in European wild boars, Bornean bearded pigs, wart hogs, and red river hogs. Three additional disrupting mutations were detected in the remaining exons. Furthermore, the rate of nonsynonymous substitutions was clearly elevated in the pig sequence compared with the corresponding sequences in humans, cattle, and mice, and we used this increased rate to estimate that UCP1 was disrupted about 20 million years ago.

Animals↗

Deletion mapping of homoeologous group 6-specific wheat expressed sequence tags.

To localize wheat (Triticum aestivum L.) ESTs on chromosomes, 882 homoeologous group 6-specific ESTs were identified by physically mapping 7965 singletons from 37 cDNA libraries on 146 chromosome, arm, and sub-arm aneuploid and deletion stocks. The 882 ESTs were physically mapped to 25 regions (bins) flanked by 23 deletion breakpoints. Of the 5154 restriction fragments detected by 882 ESTs, 2043 (loci) were localized to group 6 chromosomes and 806 were mapped on other chromosome groups. The number of loci mapped was greatest on chromosome 6B and least on 6D. The 264 ESTs that detected orthologous loci on all three homoeologs using one restriction enzyme were used to construct a consensus physical map. The physical distribution of ESTs was uneven on chromosomes with a tendency toward higher densities in the distal halves of chromosome arms. About 43% of the wheat group 6 ESTs identified rice homologs upon comparisons of genome sequences. Fifty-eight percent of these ESTs were present on rice chromosome 2 and the remaining were on other rice chromosomes. Even within the group 6 bins, rice chromosomal blocks identified by 1-6 wheat ESTs were homologous to up to 11 rice chromosomes. These rice-block contigs were used to resolve the order of wheat ESTs within each bin.

Chromosome Mapping↗

Determination and analysis of the pre-mRNA cleavage sites in Arabidopsis.

Alignment of the Arabidopsis cDNA to genome DNA sequences revealed that approximately half of the mRNAs were found to contain 3' non-templated nucleotide addition prior to the poly(A) sequence. These findings suggest that we can more precisely determine the cleavage sites. Based on the nucleotide downstream of cleavage sites, we are able to derive a hierarchy of cleavage preferences A>>U>C>>G. Interestingly, the completely different hierarchy of preferences was derived to be U>>A>>G>C, based on the nucleotide upstream of cleavage sites. Dinucleotide AU is the most preferred composition for cleavage. The determination of the cleavage sites and systematical analysis can help in the enhancement of 3'-processing site prediction and mechanistic understanding of 3'-end processing in plant.

Arabidopsis↗