Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Inconsistencies between human genetic cytolocations and those derived using genomic sequence.

One result of the publishing of the human genome sequence is the ability to define objects through their position on the consensus sequence. While this has simplified the process of creating order maps for genes on a chromosome, it has created discrepancies between the published cytolocations of human genes, as presented through genetic references, and those locations derived computationally from the genomic sequence. For the 6,830 records with HUGO gene symbols shared between the online version of Mendelian Inheritance in Man and Ensembl, 18% of the records have a discrepancy of at least one cytogenetic band between the datasets. Discordance between data sets at this frequency would have a significant impact on the utility of datasets created by the amalgamation of numerous biological databases.

Base Sequence↗

The entire genomic sequence and cDNA expression of mouse alpha-galactosidase A.

The full-length cDNA and genomic sequences encoding mouse alpha-galactosidase A (alpha-Gal A; EC 3.2.1.22), a lysosomal galactohydrolase, were isolated and characterized. The cDNA's open reading frame encoded 419 amino acids and had 82% nucleotide (nt) and 78% amino acid identity with the human sequence, although the carboxy terminus of the mouse alpha-Gal A polypeptide was 10 amino acids shorter. The functional integrity of the mouse cDNA was demonstrated by transient expression in COS-1 cells. Northern analysis revealed two mRNA species of about 1.6 and 3.4 kb due to alternative polyadenylation signals. The entire 14.4-kb mouse genomic sequence was determined; each of its seven exons was interrupted by intronic sequence at the identical positions as the exons in the human gene. The mouse 5' flanking region (250 nt) had one Sp1, site, five CAAT boxes, and no TATA box and had 67% identity with the human promoter region. The gene contained 18 complete or partial Alu-repetitive elements (13 type 1 and 5 type 2 repeats), and three putative functional AATAAA consensus polyadenylation signals were identified 72, 1668, and 1682 nt after the TAA termination codon. Use of the 72-nt site and the 1866 and/or 1682 sites were consistent with the shorter and longer transcripts. The availability of the full-length cDNA and genomic sequence encoding mouse alpha-Gal A should facilitate structure/function studies of this lysosomal glycosidase and the construction of alpha-Gal A-deficient mice by targeted gene disruption.

Amino Acid Sequence↗

Genome sequence of the non-pathogenic strain 15 of pneumonia virus of mice and comparison with the genome of the pathogenic strain J3666.

Pneumonia virus of mice (PVM) is a member of the subfamily Pneumovirinae and is the closest known relative of respiratory syncytial virus. Both viruses cause pneumonia in their respective hosts. Here, the genome sequences of two strains of PVM, non-pathogenic strain 15 and pathogenic strain J3666, are reported. Comparison of the genome sequences revealed 59 nucleotide differences between the two strains, 37 of which were coding. The nucleotide differences were spread throughout the genome, affecting cis-acting regulatory regions and seven of the ten genes. Development of a reverse-genetics system for PVM should allow further elucidation of the functional importance of the genetic differences between the two strains identified here.

Amino Acid Sequence↗

Inferring gene structures in genomic sequences using pattern recognition and expressed sequence tags.

Computational methods for gene identification in genomic sequences typically have two phases: coding region prediction and gene parsing. While there are many effective methods for predicting coding regions (exons), parsing the predicted exons into proper gene structures, to a large extent, remains an unsolved problem. This paper presents an algorithm for inferring gene structures from predicted exon candidates, based on Expressed Sequence Tags (ESTs) and biological intuition/rules. The algorithm first finds all the related ESTs in the EST database (dbEST) for each predicted exon, and infers the boundaries of one or a series of genes based on the available EST information and biological rules. Then it constructs gene models within each pair of gene boundaries, that are most consistent with the EST information. By exploiting EST information and biological rules, the algorithm can (1) model complicated multiple gene structures, including embedded genes, (2) identify falsely-predicted exons and locate missed exons, and (3) make more accurate exon boundary predictions. The algorithm has been implemented and tested on long genomic sequences with a number of genes. Test results show that very accurate (predicted) gene models can be expected when related ESTs exist for the predicted exons.

Algorithms↗

Comparisons of the genomic sequences of erysimum latent virus and other tymoviruses: a search for the molecular basis of their host specificities.

The nucleotide sequence of the genome of erysimum latent tymovirus (ELV) has been determined. It closely resembles those of the other four sequenced tymoviral genomes in its gene organization and composition, but is the smallest (6034 nucleotides) and most distinct of them. Furthermore the 78 non-coding nucleotides at the 3' terminus of the ELV genome are unable to form a complete tRNA-like structure like that reported for other tymoviruses. Comparisons of the five tymovirus genomes and their encoded proteins indicate that they have probably evolved from the progenitor tymovirus by independent progressive mutational change without genetic recombination. Comparisons of the sequences of the two non-virion proteins of five tymoviruses, and virion proteins of 17 tymoviruses, revealed no specific similarities between those of ELV and turnip yellow mosaic virus that could explain why their host ranges and symptoms are so similar, yet differ, in this respect, from ononis yellow mosaic, kennedya yellow mosaic and eggplant mosaic tymoviruses.

Amino Acid Sequence↗

Haemophilus influence: the impact of whole genome sequencing on microbiology.

The publication of the Haemophilus influenzae genome sequence in 1995 was a landmark in microbiological research. It has changed our understanding of the prokaryotic world, and will influence the approach and focus of research on microorganisms over the next few years. In this article we outline what has been learned from this and other genome sequencing projects, and discuss some of the potential avenues of investigation that will follow in the 'post-genome era'.

Biological Evolution↗

Genomic sequence analyses of segments 1 to 6 of Dendrolimus punctatus cytoplasmic polyhedrosis virus.

The complete nucleotide sequences of genomic segments S1 to S6 from Dendrolimus punctatus cypovirus 1 (DpCPV-1) have been determined. Each segment of S1 to S6 possess a single open reading frame. Conserved motifs 5' (AGUAA) and 3'(GUUAGCC) were found at the ends of each segment. Comparison of the proteins of DpCPV with those of other members in the family Reoviridae lead us to suggest that S1, S3, S4 and S6 encode the viral structural protein VP1, VP2, VP3 and VP4, respectively. S5 encoded viral non-structural protein p100 and S2 encodes an RNA-dependent RNA polymerase (RdRp). Motif analysis shows that VP3 is similar to the methyltransferase of Methanosarcina mazei Goe1, VP4 has motifs for leucine zipper and ATP/GTP-binding sites, and p100 is remarkably similar to foot-and-mouth disease virus 2A protease (FMDV 2Apro). Phylogenetic analysis of RdRps from nine viruses of the family Reoviridae indicates that DpCPV is a type 1 cypovirus, more related to Bombyx mori cypovirus (BmCPV) than to other cypovirus species. DpCPV is more related to Rice ragged stunt virus (RRSV) than to other members of different genera of the family Reoviridae, which seems to confirm the previous hypothesis that plant reoviruses originated from insect reoviruses.

Amino Acid Sequence↗

Combined evidence annotation of transposable elements in genome sequences.

Transposable elements (TEs) are mobile, repetitive sequences that make up significant fractions of metazoan genomes. Despite their near ubiquity and importance in genome and chromosome biology, most efforts to annotate TEs in genome sequences rely on the results of a single computational program, RepeatMasker. In contrast, recent advances in gene annotation indicate that high-quality gene models can be produced from combining multiple independent sources of computational evidence. To elevate the quality of TE annotations to a level comparable to that of gene models, we have developed a combined evidence-model TE annotation pipeline, analogous to systems used for gene annotation, by integrating results from multiple homology-based and de novo TE identification methods. As proof of principle, we have annotated "TE models" in Drosophila melanogaster Release 4 genomic sequences using the combined computational evidence derived from RepeatMasker, BLASTER, TBLASTX, all-by-all BLASTN, RECON, TE-HMM and the previous Release 3.1 annotation. Our system is designed for use with the Apollo genome annotation tool, allowing automatic results to be curated manually to produce reliable annotations. The euchromatic TE fraction of D. melanogaster is now estimated at 5.3% (cf. 3.86% in Release 3.1), and we found a substantially higher number of TEs (n = 6,013) than previously identified (n = 1,572). Most of the new TEs derive from small fragments of a few hundred nucleotides long and highly abundant families not previously annotated (e.g., INE-1). We also estimated that 518 TE copies (8.6%) are inserted into at least one other TE, forming a nest of elements. The pipeline allows rapid and thorough annotation of even the most complex TE models, including highly deleted and/or nested elements such as those often found in heterochromatic sequences. Our pipeline can be easily adapted to other genome sequences, such as those of the D. melanogaster heterochromatin or other species in the genus Drosophila.

Journal Article↗

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246 K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246 K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86 K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans↗

Tracking adaptive evolutionary events in genomic sequences.

As more gene and genomic sequences from an increasing assortment of species become available, new pictures of evolution are emerging. Improved methods can pinpoint where positive and negative selection act in individual codons in specific genes on specific branches of phylogenetic trees. Positive selection appears to be important in the interaction between genotype, protein structure, function, and organismal phenotype.

Animals↗

Probing meiosis in hybrids of Lolium (Poaceae) with a discriminatory repetitive genomic sequence.

A moderately repetitive genomic DNA sequence (designated pLPBB2-123) derived from Lolium perenne L. (Poaceae) is considerably more abundant in the genome of this species than in that of the closely related L. temulentum. The repetitive sequence probe is clearly able to discriminate between the genomic DNA of both species in Southern analysis, and effectively 'paints' only the chromosome set of L. perenne in diploid and triploid hybrids with L. temulentum. Fluorescence in situ hybridisation of this sequence onto homologous chromosomes during meiosis I of the hybrids shows that the sequence is evenly distributed along all of the chromosomes of L. perenne and appears to have little effect on the structural integrity or recombination potential of hybrid bivalents. Discrimination between chromatin of different parental origin in hybrid bivalents shows for the first time a progressive relaxation of relational coiling of homoeologues throughout meiotic prophase. It also highlights structural irregularities that can now be unequivocally assigned to the longer chromosomes of L. temulentum. The advantages of the use of specific differentially amplified sequences instead of whole genome probes are discussed within the context of introgression breeding programmes within the Lolium/Festuca complex.

Chromosomes↗

The NS5 gene location of two turkey meningoencephalitis virus genomic sequences.

Two new turkey meningoencephalitis virus (TMEV) nucleotide sequences were aligned to complete sequences of genomes of the flaviviruses that were available at present in the GeneBank. It was found that the both TMEV sequences represent different NS5 locations; the sequence with Acc. No. AF098456 is located downstream of that with Acc. No. AF013377 on the TMEV NS5 gene. This finding provides further insight into the TMEV NS5 gene structure and shows that the two sequences are located on the NS5 gene separately.

Animals↗

Assessing the impact of comparative genomic sequence data on the functional annotation of the Drosophila genome.

BACKGROUND: It is widely accepted that comparative sequence data can aid the functional annotation of genome sequences; however, the most informative species and features of genome evolution for comparison remain to be determined. RESULTS: We analyzed conservation in eight genomic regions (apterous, even-skipped, fushi tarazu, twist, and Rhodopsins 1, 2, 3 and 4) from four Drosophila species (D. erecta, D. pseudoobscura, D. willistoni, and D. littoralis) covering more than 500 kb of the D. melanogaster genome. All D. melanogaster genes (and 78-82% of coding exons) identified in divergent species such as D. pseudoobscura show evidence of functional constraint. Addition of a third species can reveal functional constraint in otherwise non-significant pairwise exon comparisons. Microsynteny is largely conserved, with rearrangement breakpoints, novel transposable element insertions, and gene transpositions occurring in similar numbers. Rates of amino-acid substitution are higher in uncharacterized genes relative to genes that have previously been studied. Conserved non-coding sequences (CNCSs) tend to be spatially clustered with conserved spacing between CNCSs, and clusters of CNCSs can be used to predict enhancer sequences. CONCLUSIONS: Our results provide the basis for choosing species whose genome sequences would be most useful in aiding the functional annotation of coding and cis-regulatory sequences in Drosophila. Furthermore, this work shows how decoding the spatial organization of conserved sequences, such as the clustering of CNCSs, can complement efforts to annotate eukaryotic genomes on the basis of sequence conservation alone.

Animals↗

Research needs for human nutrition in the post-genome-sequencing era.

The sequencing and annotation of the human genome and the genomes of other model organisms offer new tools and new opportunities for human nutrition research in the 21st century. Basic research continues to be the key foundation for formulating solid nutrition recommendations for the public, but the basis for establishing human nutrient recommendations today suffers because of lack of good biomarkers and because of weak federal funding for nutrition research. In the context of this post-human genome-sequencing era, tantalizing opportunity exists in seven areas-four basic and three applied: 1) identification of molecular biomarkers for nutrient status; 2) characterization of single polynuclear polymorphisms (SNPs) associated with nutrition; 3) development of a national genome array nutrition database; 4) use of models at all phylogenetic levels for nutrition research; 5) application of post-genome-sequencing tools to study diet and human health, including phytonutrients and genetically modified foods; 6) application of these tools to the role of nutrition in pathogenesis of human disease; and 7) development of a funding base for research on exercise and human health.

Biomarkers↗

Statistical properties of open reading frames in complete genome sequences.

Some statistical properties of open reading frames in all currently available complete genome sequences are analyzed (seventeen prokatyotic genomes, and 16 chromosome sequences from the yeast genome). The size distribution of open reading frames is characterized by various techniques, such as quantile tables, QQ-plots, rank-size plots (Zipf's plots), and spatial densities. The issue of the influence of CG% on the size distribution is addressed. When yeast chromosomes are compared with archaeal and eubacterial genomes, they tend to have more long open reading frames. There is little or no evidence to reject the null hypothesis that open reading frames on six different reading frames and two strands distribute similarly. A topic of current interest, the base composition asymmetry in open reading frames between the two strands, is studied using regression analysis. The base composition asymmetry at three codon positions is analyzed separately. It was shown in these genome sequences that the first codon position is G- and A-rich (i.e. purine-rich); there is a co-existence of A- and T-rich branches at the second codon position; and the third codon position is weakly T-rich.

Base Sequence↗

Whole-genome sequencing of staphylococcus haemolyticus uncovers the extreme plasticity of its genome and the evolution of human-colonizing staphylococcal species.

Staphylococcus haemolyticus is an opportunistic bacterial pathogen that colonizes human skin and is remarkable for its highly antibiotic-resistant phenotype. We determined the complete genome sequence of S.haemolyticus to better understand its pathogenicity and evolutionary relatedness to the other staphylococcal species. A large proportion of the open reading frames in the genomes of S.haemolyticus, Staphylococcus aureus, and Staphylococcus epidermidis were conserved in their sequence and order on the chromosome. We identified a region of the bacterial chromosome just downstream of the origin of replication that showed little homology among the species but was conserved among strains within a species. This novel region, designated the "oriC environ," likely contributes to the evolution and differentiation of the staphylococcal species, since it was enriched for species-specific nonessential genes that contribute to the biological features of each staphylococcal species. A comparative analysis of the genomes of S.haemolyticus, S.aureus, and S.epidermidis elucidated differences in their biological and genetic characteristics and pathogenic potentials. We identified as many as 82 insertion sequences in the S.haemolyticus chromosome that probably mediated frequent genomic rearrangements, resulting in phenotypic diversification of the strain. Such rearrangements could have brought genomic plasticity to this species and contributed to its acquisition of antibiotic resistance.

Biological Evolution↗

Genomic sequence and organization of two members of a human lectin gene family.

We have isolated and sequenced the genomic DNA encoding a human dimeric soluble lactose-binding lectin. The gene has four exons, and its upstream region contains sequences that suggest control by glucocorticoids, heat (environmental) shock, metals, and other factors. We have also isolated and sequenced three exons of the gene encoding another human putative lectin, the existence of which was first indicated by isolation of its cDNA. Comparisons suggest a general pattern of genomic organization of members of this lectin gene family.

Amino Acid Sequence↗