Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

[Mapping and human genome sequence program].

Until recently, human genome programs focused primarily on establishing maps that would provide signposts to researchers seeking to identify genes responsible for inherited diseases, as well as a basis for genome sequencing studies. Preestablished gene mapping goals have been reached. The over 7,000 microsatellite markers identified to date provide a map of sufficient density to allow localization of the gene of a monogenic disease with a precision of 1 to 2 million base pairs. The physical map, based on systematically arranged overlapping sets of artificial yeast chromosomes (YACs), has also made considerable headway during the last few years. The most recently published map covers more than 90% of the genome. However, currently available physical maps cannot be used for sequencing studies because multiple rearrangements occur in YACs. The recently developed sets of radioinduced hybrids are extremely useful for incorporating genes into existing maps. A network of American and European laboratories has successfully used these radioinduced hybrids to map 15,000 gene tags from large-scale cDNA library sequencing programs. There are increasingly pressing reasons for initiating large scale human genome sequencing studies.

Chromosome Mapping↗

Molecular variability analysis of five new complete cacao swollen shoot virus genomic sequences.

Cacao swollen shoot virus (CSSV), a member of the family Caulimovi-ridae, genus Badnavirus occurs in all the main cacao-growing areas of West Africa. We amplified, cloned and sequenced complete genomes of five new isolates, two originating from Togo and three originating from Ghana. The genome of these five newly sequenced isolates all contain the five putative open reading frames I, II, III, X and Y described for the first sequenced CSSV isolate, Agou1 originating from Togo. Their genomes have been aligned with the genome of Agou1. The nucleotide and amino acid sequence identities between isolates have been calculated and a phylogenetic analysis has been made including other pararetroviruses. Maximum nucleotide sequence variability between complete genomes of CSSV isolates was 29.4%. Geographical differentiation between isolates appears more important than differentiation between mild and severe isolates. ORF X differs greatly in size and sequence between the Togolese isolates Nyongbo2 and Agou1, and the four other isolates, its functional role is therefore clearly questionable.

Badnavirus↗

Exon discovery by genomic sequence alignment.

MOTIVATION: During evolution, functional regions in genomic sequences tend to be more highly conserved than randomly mutating 'junk DNA' so local sequence similarity often indicates biological functionality. This fact can be used to identify functional elements in large eukaryotic DNA sequences by cross-species sequence comparison. In recent years, several gene-prediction methods have been proposed that work by comparing anonymous genomic sequences, for example from human and mouse. The main advantage of these methods is that they are based on simple and generally applicable measures of (local) sequence similarity; unlike standard gene-finding approaches they do not depend on species-specific training data or on the presence of cognate genes in data bases. As all comparative sequence-analysis methods, the new comparative gene-finding approaches critically rely on the quality of the underlying sequence alignments. RESULTS: Herein, we describe a new implementation of the sequence-alignment program DIALIGN that has been developed for alignment of large genomic sequences. We compare our method to the alignment programs PipMaker, WABA and BLAST and we show that local similarities identified by these programs are highly correlated to protein-coding regions. In our test runs, PipMaker was the most sensitive method while DIALIGN was most specific. AVAILABILITY: The program is downloadable from the DIALIGN home page at http://bibiserv.techfak.uni-bielefeld.de/dialign/.

Animals↗

Rice genomics: current status of genome sequencing.

Since its establishment in 1991, the Rice Genome Research Program (RGP) has produced some basic tools for rice genome analysis, including a cDNA catalogue, a genetic linkage map and a yeast artificial chromosome (YAC)-based physical map. For the further development of rice genomics, RGP launched in 1998 an international collaborative project on rice genome sequencing. A P1-derived artificial chromosome (PAC)-based, sequence-ready physical map has been constructed using the PCR markers from cDNA sequences (expressed sequence tag [EST] markers). Selected PAC clones with 100-150 kb inserts from chromosomes 1 and 6 have been subjected to shotgun sequencing. The assembled genomic sequences, after predicting the gene-coding region, have been published both through a public database and through our website. As of January 2000, 1.9 Mb from 13 PAC clones were published. Future prospects for understanding rice genomic information at the nucleotide level are discussed.

DNA, Plant↗

Drosophila melanogaster: a case study of a model genomic sequence and its consequences.

The sequencing and annotation of the Drosophila melanogaster genome, first published in 2000 through collaboration between Celera Genomics and the Drosophila Genome Projects, has provided a number of important contributions to genome research. By demonstrating the utility of methods such as whole-genome shotgun sequencing and genome annotation by a community "jamboree," the Drosophila genome established the precedents for the current paradigm used by most genome projects. Subsequent releases of the initial genome sequence have been improved by the Berkeley Drosophila Genome Project and annotated by FlyBase, the Drosophila community database, providing one of the highest-quality genome sequences and annotations for any organism. We discuss the impact of the growing number of genome sequences now available in the genus on current Drosophila research, and some of the biological questions that these resources will enable to be solved in the future.

Animals↗

Genome sequence of Streptococcus agalactiae, a pathogen causing invasive neonatal disease.

Streptococcus agalactiae is a commensal bacterium colonizing the intestinal tract of a significant proportion of the human population. However, it is also a pathogen which is the leading cause of invasive infections in neonates and causes septicaemia, meningitis and pneumonia. We sequenced the genome of the serogroup III strain NEM316, responsible for a fatal case of septicaemia. The genome is 2 211 485 base pairs long and contains 2118 protein coding genes. Fifty-five per cent of the predicted genes have an ortholog in the Streptococcus pyogenes genome, representing a conserved backbone between these two streptococci. Among the genes in S. agalactiae that lack an ortholog in S. pyogenes, 50% are clustered within 14 islands. These islands contain known and putative virulence genes, mostly encoding surface proteins as well as a number of genes related to mobile elements. Some of these islands could therefore be considered as pathogenicity islands. Compared with other pathogenic streptococci, S. agalactiae shows the unique feature that pathogenicity islands may have an important role in virulence acquisition and in genetic diversity.

Amino Acid Sequence↗

Using fuzzy logic to confirm the integrity of a pattern recognition algorithm for long genomic sequences: the W-curve.

The W-curve is a numerical mapping algorithm that provides tertiary information content of long and short genomic sequences. The most popular genomic pattern recognition algorithms depend on string matching of the primary information content of short genomic sequences. Herein, we describe a way to define the fuzzy properties of the W-curve. This approach improves a distance (dissimilarity) between two or more homologous long genomic sequences. Fourier analysis of W-curves delivers a smoother function for gap-stripped regions. Calculation of respective Fourier energies may improve the accuracy of the distance metric used to generate a phylogenetic tree of analyzed genomic sequences. This is especially the case for long genomic sequences that have been gap-stripped and aligned with the aid of previously published heuristic methods. These previous methods involved W-curve alignments used in concert with such programs as Clustal that use linear dynamic programming to align multiple gap-stripped W-curves.

Algorithms↗

Human herpesvirus 6B genome sequence: coding content and comparison with human herpesvirus 6A.

Human herpesvirus 6 variants A and B (HHV-6A and HHV-6B) are closely related viruses that can be readily distinguished by comparison of restriction endonuclease profiles and nucleotide sequences. The viruses are similar with respect to genomic and genetic organization, and their genomes cross-hybridize extensively, but they differ in biological and epidemiologic features. Differences include infectivity of T-cell lines, patterns of reactivity with monoclonal antibodies, and disease associations. Here we report the complete genome sequence of HHV-6B strain Z29 [HHV-6B(Z29)], describe its genetic content, and present an analysis of the relationships between HHV-6A and HHV-6B. As sequenced, the HHV-6B(Z29) genome is 162,114 bp long and is composed of a 144,528-bp unique segment (U) bracketed by 8,793-bp direct repeats (DR). The genomic sequence allows prediction of a total of 119 unique open reading frames (ORFs), 9 of which are present only in HHV-6B. Splicing is predicted in 11 genes, resulting in the 119 ORFs composing 97 unique genes. The overall nucleotide sequence identity between HHV-6A and HHV-6B is 90%. The most divergent regions are DR and the right end of U, spanning ORFs U86 to U100. These regions have 85 and 72% nucleotide sequence identity, respectively. The amino acid sequences of 13 of the 17 ORFs at the right end of U differ by more than 10%, with the notable exception of U94, the adeno-associated virus type 2 rep homolog, which differs by only 2.4%. This region also includes putative cis-acting sequences that are likely to be involved in transcriptional regulation of the major immediate-early locus. The catalog of variant-specific genetic differences resulting from our comparison of the genome sequences adds support to previous data indicating that HHV-6A and HHV-6B are distinct herpesvirus species.

Base Sequence↗

The genome sequence of Rickettsia felis identifies the first putative conjugative plasmid in an obligate intracellular parasite.

We sequenced the genome of Rickettsia felis, a flea-associated obligate intracellular alpha-proteobacterium causing spotted fever in humans. Besides a circular chromosome of 1,485,148 bp, R. felis exhibits the first putative conjugative plasmid identified among obligate intracellular bacteria. This plasmid is found in a short (39,263 bp) and a long (62,829 bp) form. R. felis contrasts with previously sequenced Rickettsia in terms of many other features, including a number of transposases, several chromosomal toxin-antitoxin genes, many more spoT genes, and a very large number of ankyrin- and tetratricopeptide-motif-containing genes. Host-invasion-related genes for patatin and RickA were found. Several phenotypes predicted from genome analysis were experimentally tested: conjugative pili and mating were observed, as well as beta-lactamase activity, actin-polymerization-driven mobility, and hemolytic properties. Our study demonstrates that complete genome sequencing is the fastest approach to reveal phenotypic characters of recently cultured obligate intracellular bacteria.

Acclimatization↗

The human genome sequence: impact on health care.

The recent sequencing of the human genome, resulting from two independent global efforts, is poised to revolutionize all aspects of human health. This landmark achievement has also vindicated two differeint methodologies that can now be used to target other important large genomes. The human genome sequence has revealed several novel/surprising features notably the probable presence of a mere 30-35,000 genes. In depth comparisons have led to classification of protein families and identification of several orthologues and paralogues. Information regarding non-protein coding genes as well as regulatory regions has thrown up several new areas of research. Although still incomplete, the sequence is poised to become a boon to pharmaceutical companies with the promise of delivering several new drug targets. Several ethical concerns have also been raised and need to be addressed in earnest. This review discusses all these aspects and dwells on the possible impact of the human genome sequence on human health, medicine and also health care delivery system.

Animals↗

What is the future of electrophoresis in large-scale genomic sequencing?

Although a finished human genome reference sequence is now available, the ability to sequence large, complex genomes remains critically important for researchers in the biological sciences, and in particular, continued human genomic sequence determination will ultimately help to realize the promise of medical care tailored to an individual's unique genetic identity. Many new technologies are being developed to decrease the costs and to dramatically increase the data acquisition rate of such sequencing projects. These new sequencing approaches include Sanger reaction-based technologies that have electrophoresis as the final separation step as well as those that use completely novel, nonelectrophoretic methods to generate sequence data. In this review, we discuss the various advances in sequencing technologies and evaluate the current limitations of novel methods that currently preclude their complete acceptance in large-scale sequencing projects. Our primary goal is to analyze and predict the continuing role of electrophoresis in large-scale DNA sequencing, both in the near and longer term.

Animals↗

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans↗

Conference report: the third Bacterial Genome Sequencing Pan-European Network conference.

The third Bacterial Genome Sequencing Pan-European Network conference, held in Engelberg, Switzerland (12-15 January 2026), brought together experts from six European countries to discuss the implementation of bacterial genome sequencing in clinical microbiology and public health. Key themes included regulatory frameworks (In Vitro Diagnostic Regulation, General Data Protection Regulation), standardization, quality control, data sharing, economic evaluation, and the integration of artificial intelligence and long-read sequencing into diagnostic workflows. Across presentations, panel discussions, and workshops, participants emphasized that successful implementation of genome sequencing requires more than technical capacity: it depends on robust validation, sustainable funding, interoperable data standards, ethical governance, and interdisciplinary collaboration. The meeting highlighted that sequencing should remain question-driven and clinically meaningful, balancing cost, turnaround time, and public health impact. Overall, the conference reinforced the need for coordinated European efforts to advance responsible, standardized, and sustainable genomic surveillance and diagnostics.

bacterial genome sequencing↗

The genome sequence of Blochmannia floridanus: comparative analysis of reduced genomes.

Bacterial symbioses are widespread among insects, probably being one of the key factors of their evolutionary success. We present the complete genome sequence of Blochmannia floridanus, the primary endosymbiont of carpenter ants. Although these ants feed on a complex diet, this symbiosis very likely has a nutritional basis: Blochmannia is able to supply nitrogen and sulfur compounds to the host while it takes advantage of the host metabolic machinery. Remarkably, these bacteria lack all known genes involved in replication initiation (dnaA, priA, and recA). The phylogenetic analysis of a set of conserved protein-coding genes shows that Bl. floridanus is phylogenetically related to Buchnera aphidicola and Wigglesworthia glossinidia, the other endosymbiotic bacteria whose complete genomes have been sequenced so far. Comparative analysis of the five known genomes from insect endosymbiotic bacteria reveals they share only 313 genes, a number that may be close to the minimum gene set necessary to sustain endosymbiotic life.

Animals↗

A comparative study of detection of p53 mutations in human breast cancer by flow cytometry, single-strand conformation polymorphism and genomic sequencing.

The accuracy of immunodetection by dual parameter flow cytometry (FCM), polymerase chain reaction-mediated single strand conformation polymorphism (PCR-SSCP) and genomic sequencing to detect p53 mutations were compared. Analysis by the last two techniques was restricted to exons 5-8. Initially, 110 breast tumours were screened for p53 expression by FCM. Seventy (64%) of tumours were immunopositive. Fifteen highly immunopositive and 15 completely immunonegative tumours were selected for further analysis by PCR-SSCP and genomic sequencing. Eleven out of 15 immunopositive tumours were found to have mutation by PCR-SSCP. Genomic sequencing confirmed the presence of mutation in 10 of these 11 immunopositive tumours. Therefore, four immunopositive tumours failed to show mutation by SSCP and five by genomic sequencing. Of the 15 immunonegative tumours, one showed mutation by both PCR-SSCP and genomic sequencing and one tumour has undergone deletion of the p53 gene. Overall, immunoreactivity correlated with both PCR-SSCP and genomic sequencing in 80% of cases (24/30), and there was 96.5% (28/29) concordance between PCR-SSCP and genomic sequencing. We conclude that there is good concordance between mutations detected by PCR-SSCP and genomic sequencing, but immunochemical detection of p53 overexpression is not an absolute indicator of p53 gene mutation.

Breast Neoplasms↗

EcoGene: a genome sequence database for Escherichia coli K-12.

The EcoGene database provides a set of gene and protein sequences derived from the genome sequence of Escherichia coli K-12. EcoGene is a source of re-annotated sequences for the SWISS-PROT and Colibri databases. EcoGene is used for genetic and physical map compilations in collaboration with the Coli Genetic Stock Center. The EcoGene12 release includes 4293 genes. EcoGene12 differs from the GenBank annotation of the complete genome sequence in several ways, including (i) the revision of 706 predicted or confirmed gene start sites, (ii) the correction or hypothetical reconstruction of 61 frame-shifts caused by either sequence error or mutation, (iii) the reconstruction of 14 protein sequences interrupted by the insertion of IS elements, and (iv) pre-dictions that 92 genes are partially deleted gene fragments. A literature survey identified 717 proteins whose N-terminal amino acids have been verified by sequencing. 12 446 cross-references to 6835 literature citations and s are provided. EcoGene is accessible at a new website: http://bmb.med.miami.edu/EcoGene/EcoWeb. Users can search and retrieve individual EcoGene GenePages or they can download large datasets for incorporation into database management systems, facilitating various genome-scale computational and functional analyses.

Databases, Factual↗

Insights into cereal genomes from two draft genome sequences of rice.

Draft genome sequences have been reported for two subspecies of rice. The drafts include the sequences of an estimated 99% of all rice genes and provide major advances in our understanding of the content and complexity of cereal genomes in general and the rice genome in particular.

Computational Biology↗

Recognition of regulatory regions in genomic sequences.

For the functional interpretation of genomic sequences, effective algorithms have to be developed that will recognize regions of specific function and thus will suggest experiments for their verification. As a first step, relevant data have to be collected in an appropriate database from which suitable training sets can be extracted. In this paper, I discuss the requirements for a database that collects information about regulatory DNA sequences and describe the structure and contents of such a database (TRANSFAC). This compiled information will serve as a basis for comprehensive analysis of sites that regulate transcription, e.g., by statistical methods. It will thus facilitate the recognition of regulatory genomic sequence information and the assignment of the corresponding regulators. Moreover, it will provide all relevant data about the regulating proteins which will allow to trace back transcriptional control cascades to their origin.

Algorithms↗