Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A new superfamily of putative NTP-binding domains encoded by genomes of small DNA and RNA viruses.

Statistically significant similarity was revealed between amino acid sequences of NTP-binding pattern-containing domains which are among the most conserved protein segments in dissimilar groups of ss and dsDNA viruses (papova-, parvo-, geminiviruses and P4 bacteriophage), and RNA viruses (picorna-, como- and nepoviruses) with small genomes. Within the aligned domains of 100-120 amino acid residues, three highly conserved sequence segments have been identified, i.e. 'A' and 'B' motifs of the NTP-binding pattern, and a third, C-terminal motif 'C', not described previously. The sequence of the 'B' motif in the proteins of the new superfamily is unusually variable, with substitutions, in some of the members, of the Asp residue conserved in other NTP-binding proteins. The 'C' motif is characterized by an invariant Asn residue preceded by a stretch of hydrophobic residues. As the new superfamily included a well studied DNA and RNA helicase, T antigen of SV40, helicase function could be tentatively assigned also to the other related viral putative NTP-binding proteins. On the other hand, the possibility of different and/or multiple functions for some of these proteins is discussed.

Biological Evolution↗

Molecular cloning of an avian anti-Müllerian hormone homologue.

Anti-Müllerian hormone (AMH) is responsible for regression of the Müllerian ducts in males during embryonic development. This peptide hormone of the transforming growth factor-beta family is also believed to play a broader role in sex determination, affecting differentiation and morphogenesis of the testes. Accordingly, in mammals, AMH is produced at much higher levels in male fetuses than in female fetuses. In contrast, in birds, both male and female embryonic gonads produce AMH at high levels, although in males it is still responsible for regression of the Müllerian ducts. Its persistent expression by the embryonic ovaries and its role in female sex determination in birds is not understood. We have cloned an avian homologue to AMH. Avian AMH cDNA encodes a 644 amino acid protein that is 42% identical to human AMH overall with increased identity at the carboxyl terminus. Similarities to human AMH include motifs of sequence identity, a conserved putative plasmin cleavage site and cysteine alignments, and similar genomic intron/exon structure. Antibodies to recombinant avian AMH cross-react with recombinant human AMH and were used to show that avian AMH is glycosylated as has been shown for the human form. The avian AMH gene is transcribed in both male and female gonads but not in liver, heart, kidney or muscle.

Amino Acid Sequence↗

A dataset of estimated heterozygous individual and carrier couple frequencies for pan-ancestry carrier screening.

The data described in this publication supported the development and evaluation of pan-ancestry reproductive carrier screening panels for autosomal recessive (AR) and X-linked (XL) conditions. Raw data included combined sets of DNA variants in 1,350 AR/XL genes obtained from the ClinVar and gnomAD databases. The dataset enabled calculations of positive yield for individuals and couples across both ancestry-specific and pan-ancestry, optimised "Goldilocks"-ranked gene panels, addressing population-specific variations in the frequencies of heterozygous individuals and carrier couples. The positive yield analysis offered a performance metric for carrier screening panels, facilitating the modeling of screening performance for panels of varying sizes and composition and providing resources for optimizing panel content to ensure equity across underrepresented genetic ancestries The dataset can support ongoing research into the equitable application of carrier screening and offers significant reuse potential for refining population genetic screening practices, validating computational models, and developing frameworks to update carrier screening panels in alignment with evolving genomic data, including in underrepresented and minority populations.

Carrier screening↗

Divergent evolutionary lines of fungal cytochrome c peroxidases belonging to the superfamily of bacterial, fungal and plant heme peroxidases.

Novel open reading frames coding for cytochrome c peroxidase (CcP) belonging to the superfamily of bacterial, fungal, and plant heme peroxidases were analyzed in the available fungal genomes. Multiple sequence alignment of 71 selected peroxidase genes revealed the presence of three conserved regions essential for their function: one on the distal and two on the proximal side of the prosthetic heme group. Conserved sequence motifs on the proximal heme side are peculiar for CcPs and are responsible for their reactivity. Phylogenetic analysis performed with the distance method as well as with the maximum likelihood method revealed the existence of three distinct subfamilies of fungal CcP and their relationship to other members of the peroxidase superfamily. These divergent CcP evolutionary lines apparently evolved from a single primordial heme peroxidase gene in parallel with the evolution of ascorbate peroxidase genes. Analyzed CcPs differ significantly in their N-terminal sequences. Only subfamily I did not exhibit a presence of any signal sequence. Subfamily II members possess a well defined signal sequence allowing processing and release into mitochondrion and also in subfamily III a signal sequence was detected. Several here analyzed peroxidase genes mainly from Candida albicans and from Rhizopus oryzae can be considered interesting for the investigation of the structure-function relationship of novel CcPs revealing differences to the well documented properties of cytochrome c peroxidase from Saccharomyces cerevisiae.

Amino Acid Motifs↗

Transcriptional regulation of the mouse PNRC2 promoter by the nuclear factor Y (NFY) and E2F1.

PNRC2 (Proline-rich Nuclear Receptor Coactivator 2) was previously identified through its interaction with SF1 (steroidogenic factor 1) and has been demonstrated to be a novel coactivator for multiple nuclear receptors. In this study, PNRC2 was found to be widely expressed in mouse tissues with a strong expression in lung, spleen, ovary, thymus, and colon. Alignment of mouse genomic sequence with mouse cDNA sequence (BC006598), using mouse genome browser, defines that PNRC2 gene, located on chromosome 4, contains 3 exons: 166 bp-exon I, 205 bp-exon II, and 1526 bp-exon III. The translational start site is located in exon III. The first two exons are not translated. The 420 bp coding sequence in exon III encodes a 140 amino acid protein. To understand the molecular mechanisms that regulate the expression of PNRC2 gene, we have cloned and characterized the 5'-flanking region of the gene. Potential transcriptional start sites were determined by 5' RACE analysis. Functional analysis of the 5' flanking region of the mPNRC2 gene by deletion mutagenesis, transient transfection and luciferase assays revealed that the -67/+53 region is the minimal promoter of the mouse PNRC2 gene in HeLa cells. Within this sequence we identified two putative binding sites (inverted CCAAT box) for the transcription factor NFY (nuclear factor Y), a factor mediating cell type-specific and cell-cycle regulated expression of genes, and one binding site for E2F1, a founding member of the E2F family that displays the properties of both an oncogene and a tumor suppressor gene. Mutating each individual CCAAT site or changing the orientation of the CAATT box led to a 5-fold decrease in PNRC2 promoter activity in transient transfection experiments. Gel shift, supershift assay, and ChIP analysis demonstrated the specific binding of NFY and E2F1 proteins to the mouse PNRC2 promoter. Transient transfections and luciferase assays further revealed that overexpression of NFY enhanced-promoter activity of PNRC2 gene in a dose-dependent manner while overexpression of E2F1 strongly repressed the activity of the PNRC2 promoter. Since most genes regulated by E2F1 or NFY play a regulatory role in the cell cycle, the finding that the PNRC2 promoter is activated by NFY and repressed by E2F1 indicates that in addition to functioning as nuclear receptor coactivator, PNRC2 may also play a role in the cell cycle.

5' Flanking Region↗

Molecular cloning of bovine CD97: an EGF-TM7 molecule expressed as isoforms.

CD97, a cell surface molecule on immune cells with potential adhesive function, is a member of the epidermal growth factor seven-span transmembrane (EGF-TM7) family. We have cloned and characterized bovine CD97 and determined its expression in various cells and tissues. Based on sequence alignment with human genomic DNA, as well as human cDNA sequences encoding various CD97 isoforms, we predict that bovine CD97 mRNA occurs in four splice variants. The encoded CD97 isoforms (800, 751, 756, and 707 amino acids in length) contain different numbers of EGF-like modules, a stalk region, and a TM7 domain with a cytoplasmic tail. RT-PCR demonstrated expression of CD97 mRNA in leukocytes and several other tissues. In each cell type, CD97 mRNA encoding the shortest isoform, CD97 (EGF 1,2,5), was detected at the highest level, which was consistent with the cell surface expression of a 90-kDa polypeptide precipitated from lysates of peripheral blood mononuclear cells.

Amino Acid Sequence↗

Hepatitis B virus genotype assignment using restriction fragment length polymorphism patterns.

Hepatitis B virus (HBV) is classified into genotypes A-F, which is important for clinical and etiological investigations. To establish a simple genotyping method, 68 full-genomic sequences and 106 S gene sequences were analyzed by the molecular evolutionary method. HBV genotyping with the S gene sequence is consistent with genetic analysis using the full-genomic sequence. After alignment of the S sequences, genotype specific regions are identified and digested by the restriction enzymes, HphI, NciI, AlwI, EarI, and NlaIV. This HBV genotyping system using restriction fragment length polymorphism (RFLP) was confirmed to be correct when the PCR products of the S gene in 23 isolates collected from various countries were digested with this method. A restriction site for EarI in genotype B was absent in spite of its presence in all the other genotypes and genotype C has no restriction site for AlwI. Only genotype E is digested with NciI, while only genotype F has a restriction site for HphI. Genotype A can be distinguished by a single restriction enzyme site for NlaIV, while genotype D digestion with this enzyme results in two products that migrates at 265 and 186 bp. This simple and accurate HBV genotyping system using RFLP is considered to be useful for research on HBV.

Base Sequence↗

Structure and organisation of the pyrimidine biosynthesis pathway genes in Lactobacillus plantarum: a PCR strategy for sequencing without cloning.

This report describes the sequence and structural organisation of the pyrimidine biosynthesis pathway genes of Lactobacillus plantarum CCM 1904. It also describes an in vitro technique based on PCR for sequencing without cloning. This new technique was developed because it was impossible to clone certain parts of the L. plantarum genomic DNA in the Escherichia coli host. L. plantarum pyr genes are organised as a 9.8-kb operon with the following order: pyrR, pyrB, pyrC, pyrAA, pyrAB, pyrD, pyrF and pyrE. There are two major differences from the pyrimidine operons of Bacillus subtilis (Quinn et al., J. Bacteriol. 266 (1991) 9113-9127; Turner et al., J. Bacteriol, 176 (1994) 3708-3722) and Bacillus caldolyticus (Ghim et al., Microbiology 140 (1994) 479-491): the absence of pyrP encoding for uracil permease, and the absence of an open reading frame named orf2, whose function is unknown. Two mutually exclusive stem-loop structures were predicted at the 5'-end of L. plantarum pyr mRNA; this operon could be regulated by transcriptional attenuation under the control of PyrR. Complementation of E. coli pyrD, pyrF and pyrE mutants was obtained with a L. plantarum genomic DNA library. Alignment of the L. plantarum Pyr proteins with other known procaryotic Pyr proteins indicates that they display highly conserved regions in Gram-positive and Gram-negative bacteria.

Base Sequence↗

Molecular cloning and sequence analysis of the mouse protein C inhibitor gene.

The gene encoding mouse protein C inhibitor (mPCI) was isolated and its nucleotide sequence determined. Alignment of the genomic sequence with that of a cDNA obtained from mouse testis revealed that the mPCI gene (like the human counterpart) is composed of five exons and four introns with highly conserved exon/intron boundaries. It encodes a pre-polypeptide of 405 amino acids, which shows 63% identity with human PCI (hPCI). The putative reactive site is identical to that of hPCI from P5 to P3', suggesting a similar protease specificity. Also the putative heparin binding sites and 'hinge' regions are highly homologous in mouse and hPCI.

Amino Acid Sequence↗

Cloning and sequence analysis of a cDNA encoding the alpha-subunit of mouse beta-N-acetylhexosaminidase and comparison with the human enzyme.

cDNAs encoding the mouse beta-N-acetylhexosaminidase alpha-subunit were isolated from a mouse testis library. The longest of these (1.7 kb) was sequenced and showed 83% similarity with the human alpha-subunit cDNA sequence. The 5' end of the coding sequence was obtained from a genomic DNA clone. Alignment of the human and mouse sequences showed that all three putative N-glycosylation sites are conserved, but that the mouse alpha-subunit has an additional site towards the C-terminus. All eight cysteines in the human sequence are conserved in the mouse. There are an additional two cysteines in the mouse alpha-subunit signal peptide. All amino acids affected in Tay-Sachs-disease mutations are conserved in the mouse.

Amino Acid Sequence↗

Crystal structure of the chi:psi sub-assembly of the Escherichia coli DNA polymerase clamp-loader complex.

The chi (chi) and psi (psi) subunits of Escherichia coli DNA polymerase III form a heterodimer that is associated with the ATP-dependent clamp-loader machinery. In E. coli, the chi:psi heterodimer serves as a bridge between the clamp-loader complex and the single-stranded DNA-binding protein. We determined the crystal structure of the chi:psi heterodimer at 2.1 A resolution. Although neither chi (147 residues) nor psi (137 residues) bind to nucleotides, the fold of each protein is similar to the folds of mononucleotide-(chi) or dinucleotide-(psi) binding proteins, without marked similarity to the structures of the clamp-loader subunits. Genes encoding chi and psi proteins are found to be readily identifiable in several bacterial genomes and sequence alignments showed that residues at the chi:psi interface are highly conserved in both proteins, suggesting that the heterodimeric interaction is of functional significance. The conservation of surface-exposed residues is restricted to the interfacial region and to just two other regions in the chi:psi complex. One of the conserved regions was found to be located on chi, distal to the psi interaction region, and we identified this as the binding site for a C-terminal segment of the single-stranded DNA-binding protein. The other region of sequence conservation is localized to an N-terminal segment of psi (26 residues) that is disordered in the crystal structure. We speculate that psi is linked to the clamp-loader complex by this flexible, but conserved, N-terminal segment, and that the chi:psi unit is linked to the single-stranded DNA-binding protein via the distal surface of chi. The base of the clamp-loader complex has an open C-shaped structure, and the shape of the chi:psi complex is suggestive of a loose docking within the crevice formed by the open faces of the delta and delta' subunits of the clamp-loader.

Adenosine Triphosphate↗

Fatty acid biosynthesis in Mycobacterium tuberculosis: lateral gene transfer, adaptive evolution, and gene duplication.

Mycobacterium tuberculosis is a high GC Gram-positive member of the actinobacteria. The mycobacterial cell wall is composed of a complex assortment of lipids and is the interface between the bacterium and its environment. The biosynthesis of fatty acids plays an essential role in the formation of cell wall components, in particular mycolic acids, which have been targeted by many of the drugs used to treat M. tuberculosis infection. M. tuberculosis has approximately 250 genes involved in fatty acid metabolism, a much higher proportion than in any other organism. In silico methods have been used to compare the genome of M. tuberculosis CDC1551 to a database of 58 complete bacterial genomes. The resulting alignments were scanned for genes specifically involved in fatty acid biosynthetic pathway I. Phylogenetic analysis of these alignments was used to investigate horizontal gene transfer, gene duplication, and adaptive evolution. It was found that of the eight gene families examined, five of the phylogenies reconstructed suggest that the actinobacteria have a closer relationship with the alpha-proteobacteria than expected. This is either due to either an ancient transfer of genes or deep paralogy and subsequent retention of the genes in unrelated lineages. Additionally, adaptive evolution and gene duplication have been an influence in the evolution of the pathway. This study provides a key insight into how M. tuberculosis has developed its unique fatty acid synthetic abilities.

Acetyl-CoA Carboxylase↗

A physical map of the Myxococcus xanthus chromosome.

A physical map of the 9.2-Mbp Myxococcus xanthus DK1622 chromosome at a resolution of 25 kbp was constructed by using a strategy that is applicable to virtually all microorganisms. Segments of the chromosome were used as hybridization probes to subdivide a yeast artificial chromosome (YAC) library into groups of linked clones. The clones were aligned by comparing their EcoRI restriction patterns. The groups of YAC clones ("contigs") were oriented and aligned with the genomic restriction map by means of common genetic and physical markers such as rare restriction sites and transposon insertions. Over 95% of the genome is represented by cloned DNA. Sixty genetic loci including > 100 genes, many of which play a role in fruiting body development, have been mapped in this way. Additional genes can now be located on the chromosome map by hybridization of their sequences to the ordered set of YAC chromosomes. The mapped genetic loci account for approximately 2% of the genome.

Base Sequence↗

A novel yeast gene product, G4p1, with a specific affinity for quadruplex nucleic acids.

G4 nucleic acids are four-stranded helical structures that are formed in vitro by nucleic acids that contain guanine tracts. These structures anneal readily under physiological conditions and are unusually stable once formed. G4 nucleic acids are thought to participate in telomere function, retroviral genome dimerization, chromosome alignment during homologue pairing, and mitotic recombination, although the in vivo demonstration of these structures in any of these situations has not yet been achieved. Here we purify and characterize an activity from yeast, G4p1, which has a high and specific affinity for G4 nucleic acids. G4p1 prefers substrates containing multiple G4 domains, has an equal affinity for parallel and antiparallel G4 structures, and binds equivalently to RNA and DNA in G4 form. The Keq for G4p1 binding to a G4 DNA oligomer is 5.0 x 10(8) M-1, under near physiological conditions. G4p1 was purified and shown to derive from a 42-kDa protein (p42). We have cloned and sequenced the gene encoding p42 and show it to encode a novel protein with a region significantly homologous to bacterial methionyl-tRNA synthetase dimerization domains. We have reconstituted the G4p1 binding activity with recombinant p42 and present evidence that G4p1 is a homodimer of p42.

Amino Acid Sequence↗

Identification of the in vitro HIV-2/SIV RNA dimerization site reveals striking differences with HIV-1.

Although their genomes cannot be aligned at the nucleotide level, the HIV-1/SIVcpz and the HIV-2/SIVsm viruses are closely related lentiviruses that contain homologous functional and structural RNA elements in their 5'-untranslated regions. In both groups, the domains containing the trans-activating region, the 5'-copy of the polyadenylation signal, and the primer binding site (PBS) are followed by a short stem-loop (SL1) containing a six-nucleotide self-complementary sequence in the loop, flanked by unpaired purines. In HIV-1, SL1 is involved in the dimerization of the viral RNA, in vitro and in vivo. Here, we tested whether SL1 has the same function in HIV-2 and SIVsm RNA. Surprisingly, we found that SL1 is neither required nor involved in the dimerization of HIV-2 and SIV RNA. We identified the NarI sequence located in the PBS as the main site of HIV-2 RNA dimerization. cis and trans complementation of point mutations indicated that this self-complementary sequence forms symmetrical intermolecular interactions in the RNA dimer and suggested that HIV-2 and SIV RNA dimerization proceeds through a kissing loop mechanism, as previously shown for HIV-1. Furthermore, annealing of tRNA(3)(Lys) to the PBS strongly inhibited in vitro RNA dimerization, indicating that, in vivo, the intermolecular interaction involving the NarI sequence must be dissociated to allow annealing of the primer tRNA.

Binding Sites↗

Triticum durum metallothionein. Isolation of the gene and structural characterization of the protein using solution scattering and molecular modeling.

A novel gene sequence, with two exons and one intron, encoding a metallothionein (MT) has been identified in durum wheat Triticum durum cv. Balcali85 genomic DNA. Multiple alignment analyses on the cDNA and the translated protein sequences showed that T. durum MT (dMT) can be classified as a type 1 MT. dMT has three Cys-X-Cys motifs in each of the N- and C-terminal domains and a 42-residue-long hinge region devoid of cysteines. dMT was overexpressed in Escherichia coli as a fusion protein (GSTdMT), and bacteria expressing the fusion protein showed increased tolerance to cadmium in the growth medium compared with controls. Purified GSTdMT was characterized by SDS- and native-PAGE, size exclusion chromatography, and matrix-assisted laser desorption ionization time-of-flight mass spectrometry. It was shown that the recombinant protein binds 4 +/- 1 mol of cadmium/mol of protein and has a high tendency to form stable oligomeric structures. The structure of GSTdMT and dMT was investigated by synchrotron x-ray solution scattering and computational methods. X-ray scattering measurements indicated a strong tendency for GSTdMT to form dimers and trimers in solution and yielded structural models that were compatible with a stable dimeric form in which dMT had an extended conformation. Results of homology modeling and ab initio solution scattering approaches produced an elongated dMT structure with a long central hinge region. The predicted model and those obtained from x-ray scattering are in agreement and suggest that dMT may be involved in functions other than metal detoxification.

Amino Acid Sequence↗

Accurate detection of very sparse sequence motifs.

Protein sequence alignments are more reliable the shorter the evolutionary distance. Here, we align distantly related proteins using many closely spaced intermediate sequences as stepping stones. Such transitive alignments can be generated between any two proteins in a connected set, whether they are direct or indirect sequence neighbors in the underlying library of pairwise alignments. We have implemented a greedy algorithm, MaxFlow, using a novel consistency score to estimate the relative likelihood of alternative paths of transitive alignment. In contrast to traditional profile models of amino acid preferences, MaxFlow models the probability that two positions are structurally equivalent and retains high information content across large distances in sequence space. Thus, MaxFlow is able to identify sparse and narrow active-site sequence signatures which are embedded in high-entropy sequence segments in the structure based multiple alignment of large diverse enzyme superfamilies. In a challenging benchmark based on the urease superfamily, MaxFlow yields better reliability and double coverage compared to available sequence alignment software. This promises to increase information returns from functional and structural genomics, where reliable sequence alignment is a bottleneck to transferring the functional or structural characterization of model proteins to entire protein superfamilies.

Actins↗

Genome wide identification and classification of alternative splicing based on EST data.

MOTIVATION: Alternative splicing is currently seen to explain the vast disparity between the number of predicted genes in the human genome and the highly diverse proteome. The mapping of expressed sequences tag (EST) consensus sequences derived from the GeneNest database onto the genome provides an efficient way of predicting exon-intron boundaries, gene structure and alternative splicing events. However, the alternative splicing events are obscured by a large number of putatively artificial exon boundaries arising due to genomic contamination or alignment errors. The current work describes a methodology to associate quality values to the predicted exon-intron boundaries. High quality exon-intron boundaries are used to predict constitutive and alternative splicing ranked by confidence values, aiming to facilitate large-scale analysis of alternative splicing and splicing in general. RESULTS: Applying the current methodology, constitutive splicing is observed in 33,270 EST clusters, out of which 45% are alternatively spliced. The classification derived from the computed confidence values for 17 of these splice events frequently correlate (15/17) with RT-PCR experiments performed for 40 different tissue samples. As an application of the confidence measure, an evaluation of distribution of alternative splicing revealed that majority of variants correspond to the coding regions of the genes. However, still a significant fraction maps to non-coding regions, thereby indicating a functional relevance of alternative splicing in untranslated regions. AVAILABILITY: The predicted alternative splice variants are visualized in the SpliceNest database at http://splicenest.molgen.mpg.de

Algorithms↗