Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Automated classification of alternative splicing and transcriptional initiation and construction of visual database of classified patterns.

MOTIVATION: Large-scale detection and classification of alternative splicing and transcriptional initiation (ASTI) is the first step towards detailed studies of the functional implication and mechanisms of these phenomena. RESULTS: We have developed an algorithm that classifies all observed units of ASTI into an extendable set of distinct types (e.g. cassette type) by converting a collection of alignments between a genomic DNA sequence and cDNA sequences into binary description. This description system can uniquely and compactly encode not only typical patterns but also any rare patterns that are usually collectively assigned to 'others.' More than 150 distinct ASTI types were found when this system was applied to genome-wide detection of ASTI units in human and five other eukaryotes. AVAILABILITY: The data detected by this system are available through ASTRA (http://alterna.cbrc.jp/), a database equipped with a Java-based browser that can interactively reorganize the order of displayed splicing patterns on demand.

Algorithms↗

Characterization of the genomic structure and tissue-specific promoter of the human nuclear receptor NR5A2 (hB1F) gene.

The human homologue of the Drosophila melanogaster orphan nuclear receptor fushi tarazu factor 1 (Ftz-F1), NR5A2 (hB1F), was initially identified as a regulatory factor that binds and activates enhancer II of hepatitis B virus. NR5A2 (hB1F) is expressed specifically in pancreas and liver, playing important roles in the regulation of several liver-specific genes. A detailed analysis on the genomic structure and promoter activity will greatly promote future studies on the function of the NR5A2 (hB1F) gene. In this report, a bacterial artificial chromosome clone and several phage clones covering the NR5A2 (hB1F) gene were isolated and the complete genomic sequence was obtained. Alignment of different cDNAs of the NR5A2 (hB1F) gene with the genomic sequence facilitated the delineation of its structural organization, which spans over 150 kb and consists of eight exons interrupted by seven introns. RT-PCR and 3'-RACE revealed that utilization of two polyadenylation signals results in the 3.8 and 5.2 kb transcripts that were observed previously. The transcription start site of the NR5A2 (hB1F) gene was mapped downstream of a canonical TATA box. An upstream fragment containing binding sites for several liver-specific and ubiquitous transcription factors exhibits hepatocyte-specific promoter activity. Transient transfections indicated that hepatocyte nuclear factors HNF1 and HNF3beta could activate NR5A2 (hB1F) promoter.

3T3 Cells↗

Rearrangement of the human tre oncogene by homologous recombination between Alu repeats of nucleotide sequences from two different chromosomes.

The rearranged region of the tre oncogene originating from chromosomes 5q23q31 and 18q12 was cloned from tumor genomic DNA, sequenced and aligned with wild-type sequences cloned from a normal human genomic library. In the breakpoint region each wild-type sequence contained two Alu repeats. The recombination occurred between the 3'-most Alu from chromosome 5 and the 5'-most Alu from chromosome 18 and, consequently, resulted in a hybrid Alu flanked with one Alu on either side. The recombinant joint was located to a 20-bp homology region in left arms of the Alu repeats involved in recombination. The same homology region was identified in the hybrid Alu of the rearranged tre. At its 5' extremity the homology region overlaps the B box of Alu-borne RNA polymerase III promoter. The 100% identity score in the region of homology suggests that the recombination process was conservative and not error prone.

Base Sequence↗

Phylogenetic analysis of the tenascin gene family: evidence of origin early in the chordate lineage.

BACKGROUND: Tenascins are a family of glycoproteins found primarily in the extracellular matrix of embryos where they help to regulate cell proliferation, adhesion and migration. In order to learn more about their origins and relationships to each other, as well as to clarify the nomenclature used to describe them, the tenascin genes of the urochordate Ciona intestinalis, the pufferfish Tetraodon nigroviridis and Takifugu rubripes and the frog Xenopus tropicalis were identified and their gene organization and predicted protein products compared with the previously characterized tenascins of amniotes. RESULTS: A single tenascin gene was identified in the genome of C. intestinalis that encodes a polypeptide with domain features common to all vertebrate tenascins. Both pufferfish genomes encode five tenascin genes: two tenascin-C paralogs, a tenascin-R with domain organization identical to mammalian and avian tenascin-R, a small tenascin-X with previously undescribed GK repeats, and a tenascin-W. Four tenascin genes corresponding to tenascin-C, tenascin-R, tenascin-X and tenascin-W were also identified in the X. tropicalis genome. Multiple sequence alignment reveals that differences in the size of tenascin-W from various vertebrate classes can be explained by duplications of specific fibronectin type III domains. The duplicated domains are encoded on single exons and contain putative integrin-binding motifs. A phylogenetic tree based on the predicted amino acid sequences of the fibrinogen-related domains demonstrates that tenascin-C and tenascin-R are the most closely related vertebrate tenascins, with the most conserved repeat and domain organization. Taking all lines of evidence together, the data show that the tenascins referred to as tenascin-Y and tenascin-N are actually members of the tenascin-X and tenascin-W gene families, respectively. CONCLUSION: The presence of a tenascin gene in urochordates but not other invertebrate phyla suggests that tenascins may be specific to chordates. Later genomic duplication events led to the appearance of four family members in vertebrates: tenascin-C, tenascin-R, tenascin-W and tenascin-X.

Animals↗

Gene and genome duplication in Acanthamoeba polyphaga Mimivirus.

Gene duplication is key to molecular evolution in all three domains of life and may be the first step in the emergence of new gene function. It is a well-recognized feature in large DNA viruses but has not been studied extensively in the largest known virus to date, the recently discovered Acanthamoeba polyphaga Mimivirus. Here, I present a systematic analysis of gene and genome duplication events in the mimivirus genome. I found that one-third of the mimivirus genes are related to at least one other gene in the mimivirus genome, either through a large segmental genome duplication event that occurred in the more remote past or through more recent gene duplication events, which often occur in tandem. This shows that gene and genome duplication played a major role in shaping the mimivirus genome. Using multiple alignments, together with remote-homology detection methods based on Hidden Markov Model comparison, I assign putative functions to some of the paralogous gene families. I suggest that a large part of the duplicated mimivirus gene families are likely to interfere with important host cell processes, such as transcription control, protein degradation, and cell regulatory processes. My findings support the view that large DNA viruses are complex evolving organisms, possibly deeply rooted within the tree of life, and oppose the paradigm that viral evolution is dominated by lateral gene acquisition, at least in regard to large DNA viruses.

Acanthamoeba↗

Conserved structural elements in glutathione transferase homologues encoded in the genome of Escherichia coli.

Multiple sequence alignments of the eight glutathione (GSH) transferase homologues encoded in the genome of Escherichia coli were used to define a consensus sequence for the proteins. The consensus sequence was analyzed in the context of the three-dimensional structure of the gst gene product (EGST) obtained from two different crystal forms of the enzyme. The enzyme consists of two domains. The N-terminal region (domain I) has a thioredoxin-like alpha/beta-fold, while the C-terminal domain (domain II) is all alpha-helical. The majority of the consensus residues (12/17) reside in the N-terminal domain. Fifteen of the 17 residues are involved in hydrophobic core interactions, turns, or electrostatic interactions between the two domains. The results suggest that all of the homologues retain a well-defined group of structural elements both in and between the N-terminal alpha/beta domain and the C-terminal domain. The conservation of two key residues for the recognition motif for the gamma-glutamyl-portion of GSH indicates that the homologues may interact with GSH or GSH analogues such as glutathionylspermidine or alpha-amino acids. The genome context of two of the homologues forms the basis for a hypothesis that the b2989 and yibF gene products are involved in glutathionylspermidine and selenium biochemistry, respectively.

Amino Acid Sequence↗

[The chromosomal location of AD7C-NTP gene and the prediction of transmembrane domains of its deduced protein].

AIM: To characterize the chromosomal location of AD7C-NTP gene and predict the transmembrane domains and sub-cellular location of its deduced protein. METHODS: The AD7C-NTP mRNA sequence was alignmented with human genomic DNA sequence by Blat server. The transmembrane domains and sub-cellular location of AD7C-NTP protein were predicted by using PHDhtm, TMHMM2.0, HMMTOP2.0, SMART and PSORT servers, etc. RESULTS: The AD7C-NTP gene located in minus strand of 1p36.11, without intron. The AD7C-NTP protein was predicted to have 3 potential transmembrane domains and locate on peroxisome's membrane. CONCLUSION: Bioinformatics analysis of the AD7C-NTP gene and its deduced protein provides valuable clues for further gene cloning and study of function.

Alzheimer Disease↗

SPAP2, an Ig family receptor containing both ITIMs and ITAMs.

This study reports cloning and characterization of SPAP2, a novel transmembrane protein. The extracellular portion of SPAP2 contains six immunoglobulin-like domains and its intracellular segment has two immunoreceptor tyrosine-based activation motifs (ITAMs) and two immunoreceptor tyrosine-based inhibition motifs (ITIMs). We also identified four alternatively spliced products. Sequence alignment with the genomic database revealed that the SPAP2 gene contains 16 exons and is localized at chromosome 1q21. PCR analyses demonstrated that SPAP2 mRNA is expressed in restricted human tissues including the kidney, salivary gland, adrenal gland, uterus, and bone marrow. Tyrosine-phosphorylated SPAP2 is specifically associated with SH2 domain-containing tyrosine kinases Syk and Zap70 and SH2 domain-containing tyrosine phosphatases SHP-1 and SHP-2. Site-specific mutagenesis studies revealed that tyrosyl residues 650 and 662 embedded in the ITIMs are responsible for the binding of Syk and Zap70 while tyrosyl residues 692 and 722 embedded in the ITIMs are involved in interactions with SHP-1 and SHP-2. Finally, recruitment of SHP-1 to the tyrosine-phosphorylated ITIMs led to a marked activation of the enzyme.

Alternative Splicing↗

Cloning and expression of secretagogin, a novel neuroendocrine- and pancreatic islet of Langerhans-specific Ca2+-binding protein.

We have cloned a novel pancreatic beta cell and neuroendocrine cell-specific calcium-binding protein termed secretagogin. The cDNA obtained by immunoscreening a human pancreatic cDNA library using the recently described murine monoclonal antibody D24 contains an open reading frame of 828 base pairs. This codes for a cytoplasmic protein with six putative EF finger hand calcium-binding motifs. The gene could be localized to chromosome 6 by alignment with GenBank genomic sequence data. Northern blot analysis demonstrated abundant expression of this protein in the pancreas and to a lesser extent in the thyroid, adrenal medulla, and cortex. In addition it was expressed in scant quantity in the gastrointestinal tract (stomach, small intestine, and colon). Thyroid tissue expression of secretagogin was restricted to C-cells. Using a sandwich capture enzyme-linked immunosorbent assay with a detection limit of 6.5 pg/ml, considerable amounts of constitutively secreted protein could be measured in tissue culture supernatants of stably transfected RIN-5F and dog insulinoma (INS-H1) cell clones; however, in stably transfected Jurkat cells, the protein was only secreted upon CD3 stimulation. Functional analysis of transfected cell lines expressing secretagogin revealed an influence on calcium flux and cell proliferation. In RIN-5F cells, the antiproliferative effect is possibly due to secretagogin-triggered down-regulation of substance P transcription.

Amino Acid Sequence↗

The halophilic archaeon Halogranum roseipondis sp. nov. is susceptible to a virus carrying an exceptionally high number of viral tRNA genes.

UNLABELLED: Archaea constitute a diverse group of organisms, many of which inhabit extreme environments, such as haloarchaea that dominate hypersaline ecosystems, like solar salterns. Sampling of solar salterns and other hypersaline environments has resulted in numerous haloarchaeal isolates, including 3 classified and 27 uncharacterized Halogranum species. However, no complete genome has so far been reported for any member of this genus. Here, we present the first comprehensive study of Halogranum sp. SS5-1 isolated from a solar saltern in Samut Sakhon, Thailand. Hgn. SS5-1 is a pleomorphic, aerobic heterotroph that thrives in high salinity and moderate temperature and is capable of hydrolyzing starch. Its genome consists of a 3.6 Mbp chromosome and seven additional plasmids. Based on our phylogenetic analyses, which establish Hgn. SS5-1 as a distinct species, we propose that it will be classified as the novel species Halogranum roseipondis sp. nov. SS5-1T. Additionally, we report that Hgn. roseipondis sp. nov. SS5-1T is infected by Hagravirus capitaneum (HGTV-1), the only virus known to infect a Halogranum host. HGTV-1 exhibits a unique head-tailed morphology and encodes the largest archaeal virus double-stranded DNA genome known to date, including 34 tRNA-encoding genes. Codon usage analysis of the viral genome suggests partial alignment with host preferences, yet the abundance of viral tRNA genes hints at broader roles, potentially including roles in translation and host regulation. This study establishes Hgn. roseipondis and HGTV-1 as a novel virus-host system, opening avenues to explore infection dynamics and the roles of virus-encoded tRNA in archaea. IMPORTANCE: Archaea that thrive in high-salinity environments are key players in geochemical cycles and important contributors to ecosystem productivity. Despite their ecological significance and importance for the development of novel methodologies in synthetic biology, haloarchaea remain poorly studied. Further exploration of haloarchaea is required to obtain valuable information on the evolution of cellular complexity and the molecular mechanisms that allow cells to thrive in harsh environmental conditions. Here, we present the characterization of a novel archaeon, Halogranum roseipondis sp. SS5-1T, alongside the infection cycle of its associated virus, Hagravirus capitaneum. This tailed myovirus carries an extraordinary set of 34 viral tRNA genes, a feature that opens intriguing questions about virus-host interactions and translational control. Our findings lay the groundwork for future investigations into the expression and function of viral tRNAs in an archaeal model system, thereby opening a new frontier for studying archaeal translation and virus-driven modulation of host cellular processes.

Halobacteriaceae↗

Evaluation of EST-data using the genome assembly.

Using expressed sequence tag (EST) data for genomewide studies requires thorough understanding of the nature of the problems that are related to handling these sequences. We investigated how EST clustering performs when the genome is used as guidance as compared to pairwise sequence alignment methods. We show that clustering with the genome as a template outperforms sequence similarity methods used to create other EST clusters, such as the UniGene set, in respect to the extent ESTs originating from the same transcriptional unit are separated into disjunct clusters. Using our approach, approximately 80% of the RefSeq genes were represented by a single EST cluster and 20% comprised of two or more EST clusters. In contrast, approximately 25% of all RefSeq genes were found to be represented by a single cluster for the UniGene clustering method. The approach minimizes the risk for overestimations due to the amount of disjunct clusters originating from the same transcript. We have also investigated the quality of EST-data by aligning ESTs to the genome. The results show how many ESTs are not adequately trimmed in respect of vector sequences and low quality regions. Moreover, we identified important problems related to ESTs aligned to the genome using BLAT, such as inferring splice junctions, and explained this aspect by simulations with synthetic data. EST-clusters created with the method are available upon request from the authors.

Base Sequence↗

Using the chicken genome sequence in the development and mapping of genetic markers in the turkey (Meleagris gallopavo).

The efficacy of employing the chicken genome sequence in developing genetic markers and in mapping the turkey genome was studied. Eighty previously uncharacterized microsatellite markers were identified for the turkey using BLAST alignment to the chicken genome. The chicken sequence was then used to develop primers for polymerase chain reaction where the turkey sequence was either unavailable or insufficient. A total of 78 primer sets were tested for amplification and polymorphism in the turkey, and informative markers were genetically mapped. Sixty-five (83%) amplified turkey genomic DNA, and 33 (42%) were polymorphic in the University of Minnesota/Nicholas Turkey Breeding Farms mapping families. All but one marker genetically mapped to the position predicted from the chicken genome sequence. These results demonstrate the usefulness of the chicken sequence for the development of genomic resources in other avian species.

Alleles↗

A new type of papillomavirus DNA, its presence in genital cancer biopsies and in cell lines derived from cervical cancer.

DNA of a new papillomavirus type was cloned from a cervical carcinoma biopsy. Two EcoRI clones of 7.8 and 6.9 kb in length were obtained, the latter contained a 900-bp deletion. The BamHI fragments of both clones were used to characterize the DNA. It represents a distinct type of papillomavirus as determined by its size, its cross-hybridization with DNA of other papillomavirus types under conditions of low stringency only, the co-linear alignment of its genome with HPV 6 and HPV 16 prototypes and its occasional occurrence as oligomeric episomes. We tentatively propose to designate it as HPV 18. DNA hybridizing with HPV 18 under stringent conditions was detected in 9/36 cervical carcinomas from Africa and Brazil, in 2/13 cervical tumors from Germany and 1/10 penile carcinomas. Benign tumors (17 cervical dysplasias, 29 genital warts), eight carcinomata in situ and 15 biopsies of normal cervical tissue were devoid of detectable HPV 18 DNA. HPV 18-related DNA was found, however, in cells of the HeLa, KB and C4-1 lines all derived from cervical cancer. The state of the viral DNA was investigated in four cervical cancer biopsies. The data reveal that the DNA might be integrated into the host cell genome. One tumor provided evidence for head to tail tandem repeats some of which persisted as circular episomes.

Base Sequence↗

Molecular cloning and characterization of SPAP1, an inhibitory receptor.

We have cloned a novel cell-surface protein designated SPAP1a for SH2 domain-containing phosphatase anchor protein 1a. SPAP1a belongs to the group of type I transmembrane proteins. Its extracellular domain contains a single immunoglobulin-like domain, and its intracellular segment has two immunoreceptor tyrosine-based inhibition motifs (ITIMs). We also identified two alternatively spliced products that were named SPAP1b and SPAP1c. SPAP1b contains a short intracellular part without ITIMs, while SPAP1c lacks the transmembrane segment and represents a potential soluble protein. Sequence alignment with the genomic database revealed that the SPAP1 gene contains seven exons and is localized at chromosome 1q21. PCR analyses demonstrated that SPAP1a mRNA is specifically expressed in human hematopoietic tissues including spleen, peripheral blood, and bone marrow, and it may be restricted to expression in B cells. Recombinant SPAP1a is tyrosine phosphorylated in cells upon pervanadate stimulation and tyrosine-phosphorylated SPAP1a recruits the SH2 domain containing phosphatase SHP-1, but not SHP-2. As a specific anchor protein of SHP-1, SPAP1a may have an important role in hematopoietic cell signaling.

Alternative Splicing↗

The structural organization of the human skeletal muscle ryanodine receptor (RYR1) gene.

The RYR1 gene encoding the Ca2+ release channel of human skeletal muscle sarcoplasmic reticulum has been cloned and exon/intron boundaries have been determined, together with a minimum of 30 bp of intron sequence flanking each splice junction. The gene contains 106 exons, of which two are alternatively spliced. The length of the gene, determined by the alignment of 16 genomic phage clones, a cosmid clone, and several long polymerase chain reaction products, is approximately 160 kb. Exons range from 15 to 813 bp, while introns range from 85 to about 16,000 bp. Analysis of the gene has confirmed published errors in the human RYR1 cDNA and confirmed the structure of two alternatively spliced exons. The numbering of the nucleotides comprising the RYR1 cDNA and the numbering of amino acids encoded by them were corrected to account for these earlier errors and omissions. Analysis of 2.4 kb of the 5' upstream sequence indicated the presence of a CCAAT box and several Sp1 binding sites between nucleotides -200 and -60 bp, flanking the proposed transcription start site at -130 bp. Several other potential transcription factor binding sites were identified throughout the 5' sequence. Knowledge of the structure of the RYR1 gene will provide an invaluable resource for the discovery of mutations in the gene that are causal of human malignant hyperthermia and central core disease.

Amino Acid Sequence↗

Human gamma-aminobutyric acid-type A receptor alpha5 subunit gene (GABRA5): characterization and structural organization of the 5' flanking region.

The gamma-aminobutyric acid-type A receptor alpha5 subunit gene (GABRA5) is widely expressed in brain and localized to the imprinted human chromosome 15q11-q13. A combination of cDNA library screening and 5' RACE analysis led to identification of three distinct mRNA isoforms of GABRA5 in human adult and fetal brain tissues, each of which differs only in the noncoding 5' UTR sequence. Alignment of the genomic and cDNA sequences of GABRA5 revealed that the mRNA isoforms resulted from three alternative first exons 1A, 1B, and 1C. Northern blot analysis showed that the expression of GABRA5 was not only tissue specific but region specific in brain. CAT reporter assays revealed promoter elements in the 5' proximity of each first exon. The GABRA5 promoter regions lacked TATA and CCAAT boxes but contained several other consensus transcriptional factor recognition sequences. These findings suggest that the differential exon 1 usage of GABRA5 arises as a consequence of alternative promoter activation.

Adult↗

A restriction map of the bacteriophage T4 genome.

We report a detailed restriction map of the bacteriophage T4 genome and the alignment of this map with the genetic map. The sites cut by the enzymes Bg/II, XhoI, KpnI, SalI, PstI, EcoRI and HindIII have been localized. Several novel approaches including two-dimensional (double restriction) electrophoretic separations were used.

Chromosome Mapping↗

Isolation and characterization of Ty1/copia-like retrotransposons in mung bean (Vigna radiata).

Two Ty1/copia-like retrotransposons, RTvr1 and RTvr2, were isolated from mung bean (Vigna radiata (L.) Wilczek) genomic DNA and are the first complete elements of this kind to be reported in this legume. Nucleotide sequence analyses revealed that both elements are AT-rich (60% and 61%, respectively) and are flanked by a target-site duplication of 5 bp. The structures of RTvr1 and RTvr2 are those of typical long terminal repeat retrotransposons. Both transposons were able to produce putative proteins with the domain order of Gag-protease-integrase-reverse transcriptase-RNase H, indicating that RTvr1 and RTvr2 belong to the Ty1/copia-like retrotransposons. Except for a 2,500-bp insertion region in RTvr2, the overall similarity between RTvr1 and RTvr2 is 92%. Dot blots showed that these two retroelements were present at a copy number of 120 per mung bean haploid genome. Multiple sequence alignments showed that the conserved motifs of the aspartic proteases, integrase, reverse transcriptase, and the RNase H in the Ty1/copia-like group all exist in RTvr1 and RTvr2.

Amino Acid Sequence↗