Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Analysis of Arabidopsis genome sequence reveals a large new gene family in plants.

A detailed analysis of the currently available Arabidopsis thaliana genomic sequence has revealed the presence of a large number of open reading frames with homology to the stigmatic self-incompatibility (S) genes of Papaver rhoeas. The products of these potential genes are all predicted to be relatively small, basic, secreted proteins with similar predicted secondary structures. We have named these potential genes SPH (S-protein homologues). Their presence appears to have been largely missed by the prediction methods currently used on the genomic sequence. Equivalent homologues could not be detected in the human, microbial, Drosophila or C. elegans genomic databases, suggesting a function specific to plants. Preliminary RT-PCR analysis indicates that at least two members of the family (SPH1, SPH8) are expressed, with expression being greatest in floral tissues. The gene family may total more than 100 members, and its discovery not only illustrates the importance of the genome sequencing efforts, but also indicates the extent of information which remains hidden after the initial trawl for potential genes.

Arabidopsis↗

PromH: Promoters identification using orthologous genomic sequences.

Accurate prediction of promoters is fundamental for understanding gene expression patterns, cell specificity and development. In the studies of conserved features of regulatory regions of orthologous genes, it was observed that major promoter functional components such as transcription start points, TATA-boxes and regulatory motifs, are significantly more conservative than the sequences around them (70-100% compared with 30-50%). To improve promoter identification accuracy, we employed these findings in a new program, PromH, created by extending the TSSW program feature set. PromH uses linear discriminant functions that take into account conservation features and nucleotide sequences of promoter regions in pairs of orthologous genes. The program was tested on two sets of pairs of orthologous, mostly human and rodent, sequences with known transcription start sites (TSS), annotated to have TATA (21 genes, 11 orthologous pairs) and TATA-less (38 genes, 19 pairs) promoters, respectively. The program correctly predicted TSS for all 21 genes of the first set with a median deviation of 2 bp from true site location. Only for two genes, was there significant (46 and 105 bp) discrepancy between predicted and annotated TSS positions. For 38 TATA-less promoters from the second set, TSS was predicted for 27 genes, in 14 cases within 10 bp distance from annotated TSS, and in 21 cases--within 100 bp distance. Despite more discrepancies between predicted and annotated TSS for genes from the second set, these results are consistent with observations of much higher occurrence of multiple TSS in TATA-less promoters. In any case, our results show that PromH identifies TSS positions significantly more accurately than any other published promoter prediction method. The PromH program is available at http://www.softberry.com/berry.phtml?topic=promh.

Animals↗

Critically unwell infants and children with mitochondrial disorders diagnosed by ultrarapid genomic sequencing.

PURPOSE: To characterize the diagnostic and clinical outcomes of a cohort of critically ill infants and children with suspected mitochondrial disorders (MD) undergoing ultrarapid genomic testing as part of a national program. METHODS: Ultrarapid genomic sequencing was performed in 454 families (genome sequencing: n = 290, exome sequencing +/- mitochondrial DNA sequencing: n = 164). In 91 individuals, MD was considered, prompting analysis using an MD virtual gene panel. These individuals were reviewed retrospectively and scored according to modified Nijmegen Mitochondrial Disease Criteria. RESULTS: A diagnosis was achieved in 47% (43/91) of individuals, 40% (17/43) of whom had an MD. Seven additional individuals in whom an MD was not suspected were diagnosed with an MD after broader analysis. Gene-agnostic analysis led to the discovery of 2 novel disease genes, with pathogenicity validated through targeted functional studies (CRLS1 and MRPL39). Functional studies enabled diagnosis in another 4 individuals. Of the 24 individuals ultimately diagnosed with an MD, 79% had a change in management, which included 53% whose care was redirected to palliation. CONCLUSION: Ultrarapid genetic diagnosis of MD in acutely unwell infants and children is critical for guiding decisions about the need for additional investigations and clinical management.

Humans↗

Identification of polymorphic tandem repeats by direct comparison of genome sequence from different bacterial strains: a web-based resource.

BACKGROUND: Polymorphic tandem repeat typing is a new generic technology which has been proved to be very efficient for bacterial pathogens such as B. anthracis, M. tuberculosis, P. aeruginosa, L. pneumophila, Y. pestis. The previously developed tandem repeats database takes advantage of the release of genome sequence data for a growing number of bacteria to facilitate the identification of tandem repeats. The development of an assay then requires the evaluation of tandem repeat polymorphism on well-selected sets of isolates. In the case of major human pathogens, such as S. aureus, more than one strain is being sequenced, so that tandem repeats most likely to be polymorphic can now be selected in silico based on genome sequence comparison. RESULTS: In addition to the previously described general Tandem Repeats Database, we have developed a tool to automatically identify tandem repeats of a different length in the genome sequence of two (or more) closely related bacterial strains. Genome comparisons are pre-computed. The results of the comparisons are parsed in a database, which can be conveniently queried over the internet according to criteria of practical value, including repeat unit length, predicted size difference, etc. Comparisons are available for 16 bacterial species, and the orthopox viruses, including the variola virus and three of its close neighbors. CONCLUSIONS: We are presenting an internet-based resource to help develop and perform tandem repeats based bacterial strain typing. The tools accessible at http://minisatellites.u-psud.fr now comprise four parts. The Tandem Repeats Database enables the identification of tandem repeats across entire genomes. The Strain Comparison Page identifies tandem repeats differing between different genome sequences from the same species. The "Blast in the Tandem Repeats Database" facilitates the search for a known tandem repeat and the prediction of amplification product sizes. The "Bacterial Genotyping Page" is a service for strain identification at the subspecies level.

Bacteria↗

ESPERR: learning strong and weak signals in genomic sequence alignments to identify functional elements.

Genomic sequence signals - such as base composition, presence of particular motifs, or evolutionary constraint - have been used effectively to identify functional elements. However, approaches based only on specific signals known to correlate with function can be quite limiting. When training data are available, application of computational learning algorithms to multispecies alignments has the potential to capture broader and more informative sequence and evolutionary patterns that better characterize a class of elements. However, effective exploitation of patterns in multispecies alignments is impeded by the vast number of possible alignment columns and by a limited understanding of which particular strings of columns may characterize a given class. We have developed a computational method, called ESPERR (evolutionary and sequence pattern extraction through reduced representations), which uses training examples to learn encodings of multispecies alignments into reduced forms tailored for the prediction of chosen classes of functional elements. ESPERR produces a greatly improved Regulatory Potential score, which can discriminate regulatory regions from neutral sites with excellent accuracy ( approximately 94%). This score captures strong signals (GC content and conservation), as well as subtler signals (with small contributions from many different alignment patterns) that characterize the regulatory elements in our training set. ESPERR is also effective for predicting other classes of functional elements, as we show for DNaseI hypersensitive sites and highly conserved regions with developmental enhancer activity. Our software, training data, and genome-wide predictions are available from our Web site (http://www.bx.psu.edu/projects/esperr).

Algorithms↗

Population genetic implications from DNA polymorphism in random human genomic sequences.

Denaturing high performance liquid chromatography (DHPLC) in combination with dye-terminator sequencing was used to survey 516 random genomic sequence tagged sites (STSs) for biallelic polymorphisms in 24 representatives of the major ethnic groups residing in the United States. Of the 301 polymorphic STSs (58.3%), 172 contained a single simple sequence polymorphism (SSP), while 78, 35, and 16 contained 2, 3, and 4-6 SSPs, respectively. Of the 541 SSPs identified, 342 (63%), 152 (28%), and 47 (9%) were transitions, transversions, and insertions or deletions, respectively. Only 21% of the STSs contained SSPs with a minor-allele frequency >20%. The nucleotide diversity estimate for random genomic sequences theta = 8.23 x 10(-4) was on average 50% higher than that for intragenic non-coding regions of the human genome ( theta = 5.52 x 10(-4). The discrepancy in Tajima's D statistic between 22 autosomal genes (D=-1.304+/-0.622, mean+/-SD) and random STSs (D=-0.27) suggests that, in the absence of significant mutation rate heterogeneity, the more negative values for genes are a consequence of directional selection rather than population growth.

Base Sequence↗

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing.

BACKGROUND: Allan-Herndon-Dudley syndrome (AHDS) is an X-linked disorder caused by pathogenic variants in the SLC16A2 gene. Although most reported variants are found in protein-coding regions or adjacent junctions, structural variations (SVs) within non-coding regions have not been previously reported. METHODS: We investigated two male siblings with severe neurodevelopmental disorders and spasticity, who had remained undiagnosed for over a decade and were negative from exome sequencing, utilizing long-read HiFi genome sequencing. We conducted a comprehensive analysis including short-tandem repeats (STRs) and SVs to identify the genetic cause in this familial case. RESULTS: While coding variant and STR analyses yielded negative results, SV analysis revealed a novel hemizygous deletion in intron 1 of the SLC16A2 gene (chrX:74,460,691 - 74,463,566; 2,876 bp), inherited from their carrier mother and shared by the siblings. Determination of the breakpoints indicates that the deletion probably resulted from Alu/Alu-mediated rearrangements between homologous AluY pairs. The deleted region is predicted to include multiple transcription factor binding sites, such as Stat2, Zic1, Zic2, and FOXD3, which are crucial for the neurodevelopmental process, as well as a regulatory element including an eQTL (rs1263181) that is implicated in the tissue-specific regulation of SLC16A2 expression, notably in skeletal muscle and thyroid tissues. CONCLUSIONS: This report, to our knowledge, is the first to describe a non-coding deletion associated with AHDS, demonstrating the potential utility of long-read sequencing for undiagnosed patients. Although interpreting variants in non-coding regions remains challenging, our study highlights this region as a high priority for future investigation and functional studies.

Humans↗

Re-annotation of the genome sequence of Mycobacterium tuberculosis H37Rv.

Original genome annotations need to be regularly updated if the information they contain is to remain accurate and relevant. Here the complete re-annotation of the genome sequence of Mycobacterium tuberculosis strain H37Rv is presented almost 4 years after the first submission. Eighty-two new protein-coding sequences (CDS) have been included and 22 of these have a predicted function. The majority were identified by manual or automated re-analysis of the genome and most of them were shorter than the 100 codon cut-off used in the initial genome analysis. The functional classification of 643 CDS has been changed based principally on recent sequence comparisons and new experimental data from the literature. More than 300 gene names and over 1000 targeted citations have been added and the lengths of 60 genes have been modified. Presently, it is possible to assign a function to 2058 proteins (52% of the 3995 proteins predicted) and only 376 putative proteins share no homology with known proteins and thus could be unique to M. tuberculosis.

Bacterial Proteins↗

Discovery of the human genome sequence in the public and private databases.

Genomes: Much heat has been generated in discussions about the key human genome sequence databases, generated by the Human Genome Project and Celera, and what specific features each offers genome researchers. Stephen W. Scherer and Joseph Cheung, who are intense users of both, offer a personal assessment of the developing contents.

Animals↗

Complete genome sequence of the entomopathogenic and metabolically versatile soil bacterium Pseudomonas entomophila.

Pseudomonas entomophila is an entomopathogenic bacterium that, upon ingestion, kills Drosophila melanogaster as well as insects from different orders. The complete sequence of the 5.9-Mb genome was determined and compared to the sequenced genomes of four Pseudomonas species. P. entomophila possesses most of the catabolic genes of the closely related strain P. putida KT2440, revealing its metabolically versatile properties and its soil lifestyle. Several features that probably contribute to its entomopathogenic properties were disclosed. Unexpectedly for an animal pathogen, P. entomophila is devoid of a type III secretion system and associated toxins but rather relies on a number of potential virulence factors such as insecticidal toxins, proteases, putative hemolysins, hydrogen cyanide and novel secondary metabolites to infect and kill insects. Genome-wide random mutagenesis revealed the major role of the two-component system GacS/GacA that regulates most of the potential virulence factors identified.

Animals↗

Using proteomics to mine genome sequences.

We present a method for mining unannotated or annotated genome sequences with proteomic data to identify open reading frames. The region of a genome coding for a protein sequence is identified by using information from the analysis of proteins and peptides with MALDI-TOF mass spectrometry. The raw genome sequence or any unassembled contigs of an organism are theoretically cleaved into a number of equal sized but overlapping fragments, and these are then translated in all six frames into a series of virtual proteins. Each virtual protein is then subjected to a theoretical enzymatic digestion. Standard proteomic sample preparation methods are used to separate, array, and digest the proteins of interest to peptides. The masses of the resulting peptides are measured using mass spectrometry and compared to the theoretical peptide masses of the virtual proteins. The region of the genome responsible for coding for a particular protein can then be identified when there are a large number of hits between peptides from the protein and peptides from the virtual protein. The method makes no assumptions about the location of a protein in a particular gene sequence or the positions or types of start and stop codons. To illustrate this approach, all 773 proteins of Pseudomonas aeruginosa contained in SWISS-PROT were used to theoretically test the method and optimize parameters. Increasing the size of the virtual proteins results in an overall improvement in the ability to detect the coding region, at the cost of decreasing the sensitivity of the method for smaller proteins. Increasing the minimum number of matching peptides, lowering the mass error tolerance, or increasing the signal-to-noise ratio of the simulated mass spectrum, improves the ability to detect coding regions. The method is further demonstrated on experimental data from Mycobacterium tuberculosis and is also shown to work with eukaryotic organisms (e.g., Homo sapiens).

Amino Acid Sequence↗

Expression patterns of predicted genes from the C. elegans genome sequence visualized by FISH in whole organisms.

More than 10 megabases of contiguous genome sequence have been submitted to the databases by the Caenorhabditis elegans Genome Sequencing Consortium. To characterize the genes predicted from the sequence, we have developed high resolution FISH for visualization of mRNA distributions in whole animals. The high resolution and sensitivity afforded by the use of directly fluorescently labelled probes and confocal imaging permitted mRNA distributions to be recorded at the cellular and subcellular level. Expression patterns were obtained for 8 out of 10 genes in an initial test set of predicted gene sequences, indicating that FISH is an effective means of characterizing predicted genes in C. elegans.

Animals↗

The complete genomic sequence of hepatitis delta virus genotype IIb prevalent in Okinawa, Japan.

The Miyako Islands, located in the southernmost part of Japan, have been reported to be endemic for hepatitis delta virus (HDV). The majority of HDV patients in this area exhibit a relatively mild course of infection that evolves into a quiescent cirrhotic condition. The entire nucleotide sequence of the Miyako isolate (L215) of HDV obtained from a cirrhotic patient infected with HDV was determined. This isolate, L215, comprises 1682 nt and encodes 213 aa of the hepatitis delta antigen. Phylogenetic analysis showed that L215 is closely related to the Taiwanese genotype IIb HDV isolate. In addition, the predicted folding structure of the antigenomic RNA substrate was different from those of the published genotype II sequences.

Base Sequence↗

Analysis of EST-driven gene annotation in human genomic sequence.

We have performed a systematic analysis of gene identification in genomic sequence by similarity search against expressed sequence tags (ESTs) to assess the suitability of this method for automated annotation of the human genome. A BLAST-based strategy was constructed to examine the potential of this approach, and was applied to test sets containing all human genomic sequences longer than 5 kb in public databases, plus 300 kb of exhaustively characterized benchmark sequence. At high stringency, 70%-90% of all annotated genes are detected by near-identity to EST sequence; >95% of ESTs aligning with well-annotated sequences overlap a gene. These ESTs provide immediate access to the corresponding cDNA clones for follow-up laboratory verification and subsequent biologic analysis. At lower stringency, up to 97% of annotated genes were identified by similarity to ESTs. The apparent false-positive rate rose to 55% of ESTs among all sequences and 20% among benchmark sequences at the lowest stringency, indicating that many genes in public database entries are unannotated. Approximately half of the alignments span multiple exons, and thus aid in the construction of gene predictions and elucidation of alternative splicing. In addition, ESTs from multiple cDNA libraries frequently cluster over genes, providing a starting point for crude expression profiles. Clone IDs may be used to form EST pairs, and particularly to extend models by associating alignments of lower stringency with high-quality alignments. These results demonstrate that EST similarity search is a practical general-purpose annotation technique that complements pattern recognition methods as a tool for gene characterization.

Base Sequence↗

Genome sequences of the honey bee pathogens Paenibacillus larvae and Ascosphaera apis.

Genome sequences offer a broad view of host-pathogen interactions at the systems biology level. With the completion of the sequence of the honey bee, interest in the relevant pathogens is heightened. Here we report the genome sequences of two of the major pathogens of honey bees, the bacterium Paenibacillus larvae (causative agent for American foulbrood disease) and the fungus Ascosphaera apis. (causative agent for chalkbrood disease). Ongoing efforts to characterize the genomes of these species can be used to understand and mitigate the effects of two important pathogens, and will provide a contrast with pathogenic, benign and freeliving relatives.

Animals↗