Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Complete nucleotide sequence, genome organization and phylogenic analysis of the canine calicivirus.

The complete genomic sequence of canine calicivirus (CaCV) isolated from feces of a dog with diarrhea was determined. The CaCV genome, a positive-sense single-stranded RNA, contained 8513 nucleotides excluding the poly(A) tail and was longer than that of any other calicivirus strain with a completely known sequence. There were three open reading frames (ORF1, nt 12-5801; ORF2, nt 5805-7880; and ORF3, nt 7877-8278). ORF1 encoded a polyprotein (calculated Mr of 214,802) which had the conserved motifs of non-structural proteins of other caliciviruses and picornaviruses. Regions containing characteristic motifs in the non-structural polyprotein of CaCV showed highest similarity with those of the species Feline calicivirus and Vesicular exanthema of swine virus in the genus Vesivirus. Phylogenic analysis indicated that CaCV formed a distinct branch within the genus. Our results strongly suggested that CaCV is a new species in the genus Vesivirus.

Animals↗

Axonemal beta heavy chain dynein DNAH9: cDNA sequence, genomic structure, and investigation of its role in primary ciliary dyskinesia.

Dyneins are multisubunit protein complexes that couple ATPase activity with conformational changes. They are involved in the cytoplasmatic movement of organelles (cytoplasmic dyneins) and the bending of cilia and flagella (axonemal dyneins). Here we present the first complete cDNA and genomic sequences of a human axonemal dynein beta heavy chain gene, DNAH9, which maps to 17p12. The 14-kb-long cDNA is divided into 69 exons spread over 390 kb. The cDNA sequence of DNAH9 was determined using a combination of methods including 5' rapid amplification of cDNA ends, RT-PCR, and cDNA library screening. RT-PCR using nasal epithelium and testis RNA revealed several alternatively spliced transcripts. The genomic structure was determined using three overlapping BACs sequenced by the Whitehead Institute/MIT Center for Genome Research. The predicted protein, of 4486 amino acids, is highly homologous to sea urchin axonemal beta heavy chain dyneins (67% identity). It consists of an N-terminal stem and a globular C-terminus containing the four P-loops that constitute the motor domain. Lack of proper ciliary and flagellar movement characterizes primary ciliary dyskinesia (PCD), a genetically heterogeneous autosomal recessive disorder with respiratory tract infections, bronchiectasis, male subfertility, and, in 50% of cases, situs inversus (Kartagener syndrome, KS). Dyneins are excellent candidate genes for PCD and KS because in over 50% of cases the ultrastructural defects of cilia are related to the dynein complex. Genotype analysis was performed in 31 PCD families with two or more affected siblings using a highly informative dinucleotide polymorphism located in intron 26 of DNAH9. Two families with concordant inheritance of DNAH9 alleles in affected individuals were observed. A mutation search was performed in these two "candidate families," but only polymorphic variants were found. In the absence of pathogenic mutations, the DNAH9 gene has been excluded as being responsible for autosomal recessive PCD in these families.

Adenosine Triphosphate↗

The Mastomys natalensis papillomavirus: nucleotide sequence, genome organization, and phylogenetic relationship of a rodent papillomavirus involved in tumorigenesis of cutaneous epithelia.

Mastomys natalensis is a rodent of African origin afflicted with a very high incidence of skin tumors (keratoacanthomas and squamous carcinomas), which are associated with a papillomavirus, M. natalensis papillomavirus (MnPV). We have determined the genomic sequence of MnPV, which has a size of 7687 bp. The genomic organization is similar to that of other papillomaviruses, with open reading frames E6, E7, E1, E2, and E4 in the early and L2 and L1 in the late region. Due to an unusually large hinge region, the transcriptional activator E2 has a size of 542 amino acids rather than 400 to 460 amino acids, as in other papillomaviruses. An open reading frame E5 coding for a small hydrophobic membrane protein is missing, as is the case for some cutaneous human papillomaviruses (HPV). This fact, together with the composition of cis-responsive elements in its long control region and phylogenetic evaluation of segments of its E6, E1, and L1 genes, indicates a relationship of MnPV to the cottontail rabbit papillomavirus and several HPV types found in lesions of cutaneous epithelia, in particular to those that are associated with epidermodysplasia verruciformis. MnPV may be a useful model system for tumorigenesis of cutaneous epithelia in humans.

Amino Acid Sequence↗

FLEXGene repository: from sequenced genomes to gene repositories for high-throughput functional biology and proteomics.

The vast amount of information generated by the human genome sequencing project and related projects has given rise to a new paradigm in experimental biology. This new paradigm invokes the experimentation and data analysis at genome-wide scales, as well as the generation of new technologies and resources that take full advantage of the available sequence information. The Institute of Proteomics at Harvard Medical School is building a comprehensive, characterized, arrayed and flexible gene repository that will allow full exploitation of the genomic information by enabling functional genomics as well as protein expression, purification and analysis at genome wide scale. The FLEXGene repository (Full Length EXpression-ready) will contain clones representing the complete set of open reading frames (ORFs) of different organisms including H. sapiens and several pathogens and model organisms. The clones are constructed using recombination-based cloning technology so that hundreds or thousands of coding regions can be transferred into any expression vector in a parallel and timely mode, allowing the broadest variety of experiments to be carried out.

Animals↗

Chromosomal mapping and zygosity check of transgenes based on flanking genome sequences determined by genomic walking.

Transgenes can affect transgenic mice via transgene expression or via the so-called positional effect. DNA sequences can be localized in chromosomes using recently established mouse genomic databases. In this study, we describe a chromosomal mapping method that uses the genomic walking technique to analyze genomic sequences that flank transgenes, in combination with mouse genome database searches. Genomic DNA was collected from two transgenic mouse lines harboring pCAGGS-based transgenes, and adaptor-ligated, enzyme restricted genomic libraries for each mouse line were constructed. Flanking sequences were determined by sequencing amplicons obtained by PCR amplification of genomic libraries with transgene-specific and adaptor primers. The insertion positions of the transgenes were located by BLAST searches of the Ensembl genome database using the flanking sequences of the transgenes, and the transgenes of the two transgenic mouse lines were mapped onto chromosomes 11 and 3. In addition, flanking sequence information was used to construct flanking primers for a zygosity check. The zygosity (homozygous transgenic, hemizygous transgenic and non-transgenic) of animals could be identified by differential band formation in PCR analyses with the flanking primers. These methods should prove useful for genetic quality control of transgenic animals, even though the mode of transgene integration and the specificity of flanking sequences needs to be taken into account.

5' Flanking Region↗

Identification by shotgun sequencing, genomic organization, and functional analysis of a fourth arylsulfatase gene (ARSF) from the Xp22.3 region.

We recently reported the isolation of two new members of the sulfatase gene family, arylsulfatase D (ARSD) and E (ARSE), located approximately 50 kb from each other in the Xp22.3 region. Mutation analysis indicated ARSE as the gene responsible for X-linked recessive chondrodysplasia punctata. Expression of the ARSE gene in COS cells resulted in a heat-labile arylsulfatase activity that was inhibited by warfarin. At the same time, we detected the presence of a 1.2-kb fragment located at approximately 60 kb from ARSD and ARSE with significant homology to these two genes, suggesting the existence of another sulfatase gene, arylsulfatase F (ARSF), in Xp22.3. We have used a combined approach of long-range genomic sequencing and screening of cDNA libraries to isolate the ARSF gene. Expression of the ARSF cDNA in COS cells resulted in a heat-labile arylsulfatase activity that is not inhibited by warfarin, supporting our hypothesis that only ARSE is specifically inhibited by warfarin and is most likely involved in warfarin embryopathy. Genomic analysis revealed that ARSF has an intron/exon organization highly similar to those of ARSD and ARSE, which is also shared by another Xp22.3 sulfatase gene, ARSC (arylsulfatase C, also known as steroid sulfatase), with the splice sites occurring at the same position in all four genes. The data obtained from sequence analysis and presented in this paper indicate that the ARSC, ARSD, ARSE, and ARSF genes are more similar to each other than to other members of the sulfatase gene family, supporting our hypothesis that they represent a subfamily of related proteins created through duplication events that occurred in an ancestral pseudoautosomal region.

Amino Acid Sequence↗

Intermediary metabolism in sea urchin: the first inferences from the genome sequence.

The genome sequence of the purple sea urchin Strongylocentrotus purpuratus recently became available. We report the results of functional annotation and initial analysis of more than 2300 proteins predicted to be involved in metabolite transport and enzymatic conversion in sea urchin. The comparison of various reconstructed biosynthetic and catabolic pathways in sea urchin to those known in other genomes suggests the overall similarity of the sea urchin metabolism to that of the vertebrates, with relatively small but non-trivial differences from both vertebrates and protostomes. There are several examples of two parallel, non-orthologous solutions for the same molecular function in sea urchin, in contrast with the other completely sequenced metazoans that tend to contain just one version of the same function. There are also genes that appear to be close phylogenetic neighbors of plant or bacterial homologs, as opposed to homologs in other Metazoa. The evolutionary and functional significance of these variations is discussed.

Amino Acids↗

Human bZIP transcription factor gene NRL: structure, genomic sequence, and fine linkage mapping at 14q11.2 and negative mutation analysis in patients with retinal degeneration.

The NRL gene encodes an evolutionarily conserved basic motif-leucine zipper transcription factor that is implicated in regulating the expression of the photoreceptor-specific gene rhodopsin. NRL is expressed in postmitotic neuronal cells and in lens during embryonic development, but exhibits a retina-specific pattern of expression in the adult. To understand regulation of NRL expression and to investigate its possible involvement in retinopathies, we have determined the complete sequence of the human NRL gene, identified a polymorphic (CA)n repeat (identical to D14S64) within the NRL-containing cosmid, and refined its location by linkage analysis. Since a locus for autosomal recessive retinitis pigmentosa (arRP) has been linked to markers at 14q11 and since mutations in rhodopsin can lead to RP, we sequenced genomic PCR products of the NRL gene and of the rhodopsin-Nrl response element from a panel of patients representing independent families with inherited retinal degeneration. The analysis did not reveal any causative mutations in this group of patients. These investigations provide the basis for delineating the DNA sequence elements that regulate NRL expression in distinct neuronal cell types and should assist in the analysis of NRL as a candidate gene for inherited diseases/syndromes affecting visual function.

Adult↗

Retention of 1.2 kbp of 'novel' genomic sequence in two European field isolates and some vaccine strains of Fowlpox virus extends open reading frame fpv241.

The emergence of variant fowlpox viruses (FWPVs) and increasing field use of recombinants against avian influenza H5N1 emphasize the need to monitor vaccines and to distinguish them from field strains. Five commercial vaccines, two laboratory viruses and two European field isolates were characterized by PCR and sequencing at 18 loci differing between attenuated FP9 and its pathogenic progenitor. PCR failed to discriminate between the viruses and sequence determination revealed no significant differences at any locus, except for a polymorphic locus encompassed by deletion 24 (9.3 kbp) in FP9. Surprisingly, 'novel' previously unreported sequence (spanning 1.2 kbp) was found in both European field isolates and three of the vaccines. It was absent from the other two vaccines, removed by a 1.2 kbp deletion identical to that surprisingly also observed in the completely sequenced genome of FPV USDA. This locus (H9) adds a potentially useful tool for discriminating between FWPV field isolates and vaccines.

Amino Acid Sequence↗

Novel PTS proteins revealed by bacterial genome sequencing: a unique fructose-specific phosphoryl transfer protein with two HPr-like domains in Haemophilus influenzae.

The completely sequenced genome of Haemophilus influenzae has been analysed for proteins of the phosphoenolpyruvate: sugar phosphotransferase system (PTS). We show that within the fructose PTS H. influenzae possesses a novel multi-domain phosphoryl transfer protein, not previously recognized, that includes two fructose-specific HPr domains fused in tandem in a single polypeptide chain.

Bacterial Proteins↗

Two distinct modes of microsatellite mutation processes: evidence from the complete genomic sequences of nine species.

We surveyed microsatellite distribution in 10 completely sequenced genomes. Using a permutation-based statistic, we assessed for all 10 genomes whether the microsatellite distribution significantly differed from expectations. Consistent with previous reports, we observed a highly significant excess of long microsatellites. Focusing on short microsatellites containing only a few repeat units, we demonstrate that this repeat class is significantly underrepresented in most genomes. This pattern was observed across different repeat types. Computer simulations indicated that neither base substitutions nor a combination of length-dependent slippage and base substitutions could explain the observed pattern of microsatellite distribution. When we introduced one additional mutation process, a length-independent slippage (indel slippage) operating at repeats with few repetitions, our computer simulations captured the observed pattern of microsatellite distribution.

Animals↗

Microbial genome sequencing.

Complete genome sequences of 30 microbial species have been determined during the past five years, and work in progress indicates that the complete sequences of more than 100 further microbial species will be available in the next two to four years. These results have revealed a tremendous amount of information on the physiology and evolution of microbial species, and should provide novel approaches to the diagnosis and treatment of infectious disease.

Anti-Infective Agents↗

Complete genome sequence of the alkaliphilic bacterium Bacillus halodurans and genomic sequence comparison with Bacillus subtilis.

The 4 202 353 bp genome of the alkaliphilic bacterium Bacillus halodurans C-125 contains 4066 predicted protein coding sequences (CDSs), 2141 (52.7%) of which have functional assignments, 1182 (29%) of which are conserved CDSs with unknown function and 743 (18. 3%) of which have no match to any protein database. Among the total CDSs, 8.8% match sequences of proteins found only in Bacillus subtilis and 66.7% are widely conserved in comparison with the proteins of various organisms, including B.subtilis. The B. halodurans genome contains 112 transposase genes, indicating that transposases have played an important evolutionary role in horizontal gene transfer and also in internal genetic rearrangement in the genome. Strain C-125 lacks some of the necessary genes for competence, such as comS, srfA and rapC, supporting the fact that competence has not been demonstrated experimentally in C-125. There is no paralog of tupA, encoding teichuronopeptide, which contributes to alkaliphily, in the C-125 genome and an ortholog of tupA cannot be found in the B.subtilis genome. Out of 11 sigma factors which belong to the extracytoplasmic function family, 10 are unique to B. halodurans, suggesting that they may have a role in the special mechanism of adaptation to an alkaline environment.

ATP-Binding Cassette Transporters↗

Sequence. Plants join the genome sequencing bandwagon.

An international consortium announced this week that it has finished the first genome sequence of a higher plant. For plant biologists, the eagerly awaited genome of this small weed, Arabidopsis thaliana, offers a window into the genetic makeup of all plants, including key crops. And it's a clear window indeed, as the six international sequencing teams on three continents have produced a genome sequence that is more complete than that of any multicellular organism which has been published to date. Through this window, they are seeing for the first time that plants may be much more complex than many biologists have imagined.

Arabidopsis↗

Detection and analysis of alternative splicing in the silkworm by aligning expressed sequence tags with the genomic sequence.

We identified 277 alternative splice forms in silkworm genes based on aligning expressed sequence tags with genomic sequences, using a transcript assembly program. A large fraction (74%) of these alternative splices are located in protein-coding regions and alter protein products, whereas only 26% are in untranslated regions. From the alternative splices located in protein-coding regions, some (43%) affect protein domains that bind various biological molecules. The vast majority of the detected alternative forms in this study appear to be novel, and potentially affect biologically meaningful control of function in silkworm genes. Our results indicate that alternative splicing in silkworm largely produces protein diversity and functional diversity, and is a widely used mechanism for regulating gene expression.

Alternative Splicing↗

ESTMAP: a system for expressed sequence tags mapping on genomic sequences.

The completion of a number of large genome sequencing projects emphasizes the importance of protein-coding gene predictions. Most of the problems associated with gene prediction are caused by the complex exon-intron structures commonly found in eukaryotic genomes. However, information from homologous sequences can significantly improve the accuracy of the prediction. In particular, expressed sequence tags (ESTs) are very useful for this purpose, since currently existing EST collections are very large. We developed an ESTMAP system, which utilizes homology searches against a database of repetitive elements using the RepeatView program and the EST Division of GenBank using the BLASTN program. ESTMAP extracts "exact" matches with EST sequences (> 95% of homology) from BLASTN output file and predicts introns in DNA comparing ESTs and a query sequence. ESTMAP is implemented as a part of the WebGene system (http://www.cnr.it/webgene).

Base Sequence↗