Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sequencing Resource”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Synaptotagmin gene content of the sequenced genomes.

BACKGROUND: Synaptotagmins exist as a large gene family in mammals. There is much interest in the function of certain family members which act crucially in the regulated synaptic vesicle exocytosis required for efficient neurotransmission. Knowledge of the functions of other family members is relatively poor and the presence of Synaptotagmin genes in plants indicates a role for the family as a whole which is wider than neurotransmission. Identification of the Synaptotagmin genes within completely sequenced genomes can provide the entire Synaptotagmin gene complement of each sequenced organism. Defining the detailed structures of all the Synaptotagmin genes and their encoded products can provide a useful resource for functional studies and a deeper understanding of the evolution of the gene family. The current rapid increase in the number of sequenced genomes from different branches of the tree of life, together with the public deposition of evolutionarily diverse transcript sequences make such studies worthwhile. RESULTS: I have compiled a detailed list of the Synaptotagmin genes of Caenorhabditis, Anopheles, Drosophila, Ciona, Danio, Fugu, Mus, Homo, Arabidopsis and Oryza by examining genomic and transcript sequences from public sequence databases together with some transcript sequences obtained by cDNA library screening and RT-PCR. I have compared all of the genes and investigated the relationship between plant Synaptotagmins and their non-Synaptotagmin counterparts. CONCLUSIONS: I have identified and compared 98 Synaptotagmin genes from 10 sequenced genomes. Detailed comparison of transcript sequences reveals abundant and complex variation in Synaptotagmin gene expression and indicates the presence of Synaptotagmin genes in all animals and land plants. Amino acid sequence comparisons indicate patterns of conservation and diversity in function. Phylogenetic analysis shows the origin of Synaptotagmins in multicellular eukaryotes and their great diversification in animals. Synaptotagmins occur in land plants and animals in combinations of 4-16 in different species. The detailed delineation of the Synaptotagmin genes presented here, will allow easier identification of Synaptotagmins in future. Since the functional roles of many of these genes are unknown, this gene collection provides a useful resource for future studies.

Alternative Splicing↗

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n = 53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans↗

Gene discovery in the auditory system using a tissue specific approach.

Molecular methods applied to the study of hereditary hearing loss in the past decade have revealed a copious number of genes representing a great diversity of cellular functions. In some instances, identification of these genes involved in deafness disorders has provided the portal to investigations of pathways not previously envisioned to underlie human hearing. Sequence analysis of a human fetal cochlear cDNA library has been employed to generate thousands of ESTs (expressed sequence tags) that provide a transcript map of the human cochlea and are a resource of candidate genes for positional cloning endeavors.

Chromosome Mapping↗

Identification and purification of proteins from germ cell-conditioned medium (GCCM).

Germ cells are known to regulate Sertoli cell and testicular function possibly through released factor(s) or via cell-cell contact. However, the identities of many of these putative biological factors are not known. The aim of this study is to present a strategy to identify and purify germ cell-derived proteins found in germ cell-conditioned medium (GCCM) at a quantity sufficient to permit protein microsequencing. The purification scheme of a novel germ cell-derived protein from GCCM designated GC-26 is presented along with several germ cell proteins using a combination of high pressure liquid chromatography (HPLC) columns. The purity of GC-26 and other germ cell proteins were confirmed by sodium dodecyl sulfate (SDS)-polyacrylamide gel electrophoresis (PAGE) and silver staining. The identities of GC-26, a 26-kDa polypeptide, and other proteins were determined by direct protein microsequencing. These partial NH2-terminal amino acid sequences were compared with the existing databases at Protein Identification Resource (PIR), GenBank, and BLAST. These analyses revealed that these proteins are unique. This strategy should be useful for the micropurification of proteins from other biological samples and/or fluids.

Amino Acid Sequence↗

41 kilobases of analyzed sequence from the pseudoautosomal and sex-determining regions of the short arm of the human Y chromosome.

Determination of 41.2 kb of Y chromosome genomic sequence has been made from a cosmid that spans the Yp pseudoautosomal boundary and includes 18.5 kb of sequence from the patient-defined sex-determining region of the Y chromosome. An AceDB database of the sequence and the analysis data have been produced as a resource for studies of the evolution and population genetics of the Y chromosome. Comparison of the 18.5 kb from the sex determining region to the sex determining region of mouse does not locate any areas of similarity outside SRY/Sry. Indeed, no coding regions other than those previously reported can be detected anywhere in the 41 kb. The Y-specific and pseudoautosomal portions of this sequence have different repeat sequence and GC contents: this may have relevance both to the events defining the pseudoautosomal boundary and to the course of sequence evolution in the absence of recombination.

Animals↗

Identification of conserved regulatory elements by comparative genome analysis.

BACKGROUND: For genes that have been successfully delineated within the human genome sequence, most regulatory sequences remain to be elucidated. The annotation and interpretation process requires additional data resources and significant improvements in computational methods for the detection of regulatory regions. One approach of growing popularity is based on the preferential conservation of functional sequences over the course of evolution by selective pressure, termed 'phylogenetic footprinting'. Mutations are more likely to be disruptive if they appear in functional sites, resulting in a measurable difference in evolution rates between functional and non-functional genomic segments. RESULTS: We have devised a flexible suite of methods for the identification and visualization of conserved transcription-factor-binding sites. The system reports those putative transcription-factor-binding sites that are both situated in conserved regions and located as pairs of sites in equivalent positions in alignments between two orthologous sequences. An underlying collection of metazoan transcription-factor-binding profiles was assembled to facilitate the study. This approach results in a significant improvement in the detection of transcription-factor-binding sites because of an increased signal-to-noise ratio, as demonstrated with two sets of promoter sequences. The method is implemented as a graphical web application, ConSite, which is at the disposal of the scientific community at http://www.phylofoot.org/. CONCLUSIONS: Phylogenetic footprinting dramatically improves the predictive selectivity of bioinformatic approaches to the analysis of promoter sequences. ConSite delivers unparalleled performance using a novel database of high-quality binding models for metazoan transcription factors. With a dynamic interface, this bioinformatics tool provides broad access to promoter analysis with phylogenetic footprinting.

Algorithms↗

Construction and characterization of a cDNA library from 4-week-old human embryo.

Development and differentiation studies of early human embryos have been severely impeded by general difficulties in obtaining suitable samples. In order to isolate and identify new genes expressed during early human development, we constructed and characterized a PCR-based cDNA library using a 4-week-old chorion-free human embryo. The constructed cDNA library contained 6.3 x 10(6) directional recombinants, and its insert size ranged from 0.4 to 1.8 kb. The cDNA library proportionally represents the mRNA population, containing beta-actin, tPA and LINE1 repetitive sequences at the expected frequencies as in other conventionally constructed and PCR-based cDNA libraries. PCR analyses of the library for specific genes have also revealed the presence of cDNAs for developmentally important genes such as CD59, MCP, Quox-1 and ZNF268. Among the 70 randomly selected cDNA clones, 53% encoded previously known genes, 26% matched with anonymous sequences, and 17% showed no sequence similarity and were designated as human early embryo-specific ESTs. These results demonstrate the sequence complexity and relatively low redundancy of our cDNA library. Furthermore, approximately 40% of those randomly analyzed clones contained full-length encoding regions. To our knowledge, this is the first description of the PCR-based cDNA library from a 4-week-old chorion-free human embryo, and the presence of novel sequences within this library makes it a valuable and unique resource for studying gene expression and regulatory mechanisms that underlie the early process of human embryogenesis.

Cloning, Molecular↗

Comparative genomics: digging for data.

Comparative genomics is a science in its infancy. It has been driven by a huge increase in freely available genome-sequence data, and the development of computer techniques to allow whole-genome sequence analyses. Other approaches, which use hybridization as a method for comparing the gene content of related organisms, are rising alongside these more bioinformatic methods. All these approaches have been pioneered using bacterial genomes because of their simplicity and the large number of complete genome sequences available. The aim of bacterial comparative genomics is to determine what genotypic differences are important for the expression of particular traits. The benefits of such studies will be a deeper understanding of these phenomena; the possibility of exposing novel drug targets, including those for antivirulence drugs; and the development of molecular techniques that reveal patients who are infected with virulent organisms so that health care resources can be allocated appropriately. With more and more genome sequences becoming available, the rise of comparative genomics continues apace.

Bacteria↗

Sequence-based alignment of sorghum chromosome 3 and rice chromosome 1 reveals extensive conservation of gene order and one major chromosomal rearrangement.

The completed rice genome sequence will accelerate progress on the identification and functional classification of biologically important genes and serve as an invaluable resource for the comparative analysis of grass genomes. In this study, methods were developed for sequence-based alignment of sorghum and rice chromosomes and for refining the sorghum genetic/physical map based on the rice genome sequence. A framework of 135 BAC contigs spanning approximately 33 Mbp was anchored to sorghum chromosome 3. A limited number of sequences were collected from 118 of the BACs and subjected to BLASTX analysis to identify putative genes and BLASTN analysis to identify sequence matches to the rice genome. Extensive conservation of gene content and order between sorghum chromosome 3 and the homeologous rice chromosome 1 was observed. One large-scale rearrangement was detected involving the inversion of an approximately 59 cM block of the short arm of sorghum chromosome 3. Several small-scale changes in gene collinearity were detected, indicating that single genes and/or small clusters of genes have moved since the divergence of sorghum and rice. Additionally, the alignment of the sorghum physical map to the rice genome sequence allowed sequence-assisted assembly of an approximately 1.6 Mbp sorghum BAC contig. This streamlined approach to high-resolution genome alignment and map building will yield important information about the relationships between rice and sorghum genes and genomic segments and ultimately enhance our understanding of cereal genome structure and evolution.

Base Sequence↗

Physical and transcript map of a 2-Mb region in Xp22.1 containing candidate genes for X-linked mental retardation and short stature.

Genetic loci for several diseases, including X-linked nonspecific mental retardation and short stature, have been mapped to Xp22.1. In spite of the recent publications of two draft sequences for the human genome, this region seems to be largely unmapped and unsequenced. Here we report an integrated physical and transcript map of approximately 2-Mb from DXS8004 to DXS365. Using sequence tagged site (STS)-content mapping and chromosome walking, we assembled a genomic clone contig of 54 BACs and one cosmid with an estimated 4.5-fold coverage of this region. The minimum tiling path consists of 23 BACs and one cosmid. Onto this contig, we mapped 30 new STSs derived from the unique end-sequences of the BACs, three expressed sequence tags, five genes, and seven CpG islands. This integrated map provides a unique resource for the positional cloning of candidate disease genes mapping to Xp22.1 and is therefore of value for the completion of the genomic sequence of this region.

Base Sequence↗

Detection of sequence polymorphisms in red junglefowl and White Leghorn ESTs.

Over 16,000 high quality expressed sequence tags (ESTs) from red junglefowl (RJ) and White Leghorn (WL) brain and testis cDNA libraries were generated. Here, we have used this resource for detection of single nucleotide polymorphisms (SNPs), and also completed full-length sequencing of 46 pairs of clones, representing the same gene from both the RJ and WL libraries. From the main set of ESTs, which were assembled using Phrap, 746 putative SNPs were identified, of which 76% were transitions and 24% were transversions. A subset of SNPs was evaluated by sequence analysis of five RJ and five WL birds. Nine of 12 SNPs were verified in this limited sample, suggesting that a majority of the putative polymorphisms documented in this study represent real SNPs. During full-length sequencing of the 46 RJ/WL clones 100 SNPs were identified, which translated to a frequency of 1.90 SNPs/1000 bp. The number of transitions and transversions were 77% and 23%, respectively, and the proportion of non-synonymous vs. synonymous SNPs was 20% and 80%, respectively. Four large insertions/deletions were identified between the RJ and WL full-length sequences, and they appear to represent different splice variants.

Animals↗

Range-wide molecular analysis of the western pond turtle (Emys marmorata): cryptic variation, isolation by distance, and their conservation implications.

We analysed phylogeography and population genetic variation across the range of the western pond turtle (Emys marmorata) using rapidly evolving mitochondrial and nuclear DNA sequence data. Nuclear DNA sequences from two unlinked introns displayed extremely low levels of variation, but phylogenetic analyses based on mtDNA recovered four well-supported and geographically coherent clades. These included a large Northern clade composed of populations from Washington south to San Luis Obispo County, California, west of the Coast Ranges; a San Joaquin Valley clade from the southern Great Central Valley; a geographically restricted Santa Barbara clade from a limited region in Santa Barbara and Ventura counties; and a Southern clade that occurs south of the Tehachapi Mountains and west of the Transverse Range south to Baja California, Mexico. An analysis of molecular variance (amova) based on regional hydrographic units revealed that populations from the Sacramento Valley north to Washington were virtually invariant, with no evidence of population substructure among northern river drainage basins. In other areas, E. marmorata contains considerable unrecognized variation, particularly in central and southern California and in northern Baja California, Mexico. Our northern clade is congruent with the distribution of the subspecies Emys marmorata marmorata (Washington-central California). However, no clade is congruent with the distribution of the southern subspecies Emys marmorata pallida from central California-Baja. Thus, recognition of the current subspecies split is not warranted, based on the available genetic evidence. Our amova and phylogenetic results, in conjunction with a growing comparative database for other codistributed aquatic taxa, confirm the occurrence of genetic breaks across the Tehachapi Mountains and Transverse Range bounding the southern end of the Great Central Valley, and point to southern California as a rich source of cryptic genetic variation.

Analysis of Variance↗

Mitochondrial DNA phylogeography of the Mesoamerican spiny-tailed lizards (Ctenosaura quinquecarinata complex): historical biogeography, species status and conservation.

Through the examination of past and present distributions of plants and animals, historical biogeographers have provided many insights on the dynamics of the massive organismal exchange between North and South America. However, relatively few phylogeographic studies have been attempted in the land bridge of Mesoamerica despite its importance to better understand the evolutionary forces influencing this biodiversity 'hotspot'. Here we use mitochondrial DNA sequence data from fresh samples and formalin-fixed museum specimens to investigate the genetic and biogeographic diversity of the threatened Mesoamerican spiny-tailed lizards of the Ctenosaura quinquecarinata complex. Species boundaries and their phylogeographic patterns are examined to better understand their disjunct distribution. Three monophyletic, allopatric lineages are established using mtDNA phylogenetic and nested clade analyses in (i) northern: México, (ii) central: Guatemala, El Salvador and Honduras, and (iii) southern: Nicaragua and Costa Rica. The average sequence divergence observed between lineages varied between 2.0% and 3.7% indicating that they do not represent a very recent split and the patterns of divergence support the recently established nomenclature of C. quinquecarinata, Ctenosaura flavidorsalis and Ctenosaura oaxacana. Considering the geological history of Mesoamerica and the observed phylogeographic patterns of these lizards, major evolutionary episodes of their radiation in Mesoamerica are postulated and are indicative of the regions' geological complexity. The implications of these findings for the historical biogeography, taxonomy and conservation of these lizards are discussed.

Animals↗

ParPEST: a pipeline for EST data analysis based on parallel computing.

BACKGROUND: Expressed Sequence Tags (ESTs) are short and error-prone DNA sequences generated from the 5' and 3' ends of randomly selected cDNA clones. They provide an important resource for comparative and functional genomic studies and, moreover, represent a reliable information for the annotation of genomic sequences. Because of the advances in biotechnologies, ESTs are daily determined in the form of large datasets. Therefore, suitable and efficient bioinformatic approaches are necessary to organize data related information content for further investigations. RESULTS: We implemented ParPEST (Parallel Processing of ESTs), a pipeline based on parallel computing for EST analysis. The results are organized in a suitable data warehouse to provide a starting point to mine expressed sequence datasets. The collected information is useful for investigations on data quality and on data information content, enriched also by a preliminary functional annotation. CONCLUSION: The pipeline presented here has been developed to perform an exhaustive and reliable analysis on EST data and to provide a curated set of information based on a relational database. Moreover, it is designed to reduce execution time of the specific steps required for a complete analysis using distributed processes and parallelized software. It is conceived to run on low requiring hardware components, to fulfill increasing demand, typical of the data used, and scalability at affordable costs.

Algorithms↗

Selecting for functional alternative splices in ESTs.

The expressed sequence tag (EST) collection in dbEST provides an extensive resource for detecting alternative splicing on a genomic scale. Using genomically aligned ESTs, a computational tool (TAP) was used to identify alternative splice patterns for 6400 known human genes from the RefSeq database. With sufficient EST coverage, one or more alternatively spliced forms could be detected for nearly all genes examined. To identify high (>95%) confidence observations of alternative splicing, splice variants were clustered on the basis of having mutually exclusive structures, and sample statistics were then applied. Through this selection, alternative splices expected at a frequency of >5% within their respective clusters were seen for only 17%-28% of genes. Although intron retention events (potentially unspliced messages) had been seen for 36% of the genes overall, the same statistical selection yielded reliable cases of intron retention for <5% of genes. For high-confidence alternative splices in the human ESTs, we also noted significantly higher rates both of cross-species conservation in mouse ESTs and of validation in the GenBank mRNA collection. We suggest quantitative analytical approaches such as these can aid in selecting useful targets for further experimental characterization and in so doing may help elucidate the mechanisms and biological implications of alternative splicing.

Alternative Splicing↗

AGRIS: Arabidopsis gene regulatory information server, an information resource of Arabidopsis cis-regulatory elements and transcription factors.

BACKGROUND: The gene regulatory information is hardwired in the promoter regions formed by cis-regulatory elements that bind specific transcription factors (TFs). Hence, establishing the architecture of plant promoters is fundamental to understanding gene expression. The determination of the regulatory circuits controlled by each TF and the identification of the cis-regulatory sequences for all genes have been identified as two of the goals of the Multinational Coordinated Arabidopsis thaliana Functional Genomics Project by the Multinational Arabidopsis Steering Committee (June 2002). RESULTS: AGRIS is an information resource of Arabidopsis promoter sequences, transcription factors and their target genes. AGRIS currently contains two databases, AtTFDB (Arabidopsis thaliana transcription factor database) and AtcisDB (Arabidopsis thaliana cis-regulatory database). AtTFDB contains information on approximately 1,400 transcription factors identified through motif searches and grouped into 34 families. AtTFDB links the sequence of the transcription factors with available mutants and, when known, with the possible genes they may regulate. AtcisDB consists of the 5' regulatory sequences of all 29,388 annotated genes with a description of the corresponding cis-regulatory elements. Users can search the databases for (i) promoter sequences, (ii) a transcription factor, (iii) a direct target genes for a specific transcription factor, or (vi) a regulatory network that consists of transcription factors and their target genes. CONCLUSION: AGRIS provides the necessary software tools on Arabidopsis transcription factors and their putative binding sites on all genes to initiate the identification of transcriptional regulatory networks in the model dicotyledoneous plant Arabidopsis thaliana. AGRIS can be accessed from http://arabidopsis.med.ohio-state.edu.

3' Untranslated Regions↗

Transcriptome analyses of human genes and applications for proteome analyses.

By utilizing recently developed full-length cDNA technologies, large-scale cDNA sequencing was carried out by several cDNA projects. Now full-length cDNA resources cover the major part of the protein-coding human genes. Comprehensive analyses of the collected full-length cDNA data revealed not only the complete sequences of thousands of novel gene transcripts but also novel alternatively spliced isoforms of hitherto identified genes. However, it was not as easy as expected to deduce their encoded amino acid sequences based solely on the full-length cDNA sequences. It was neither always the case that the longest open reading frame corresponded to the real protein coding region nor that the first ATG was the translation initiator codon. Also, proteome-wide mass-spectrometry analysis has shown that there is an unexpectedly large population of small proteins, encoded by so-called upstream open reading frames, within the cell. Since sound manual annotations by experts were still indispensable to address these problems, an international meeting to make transcriptome-wide functional annotations of cDNAs was held, namely the H-invitational. In this meeting, functional annotations were made both manually and computationally for most of the pre-existing full-length cDNAs collected from world-wide cDNA projects. The achieved integrated information for each of the cDNAs was published as a database. It was also shown that the full-length cDNA data were useful for identifying alternative splicing variants, exact transcriptional start sites of the mRNAs and the adjacent promoter regions. Rapidly accumulating genome data as well as versatile use of the transcriptome information will shortly lay a firm foundation for proteome-level understanding of human gene networks.

Alternative Splicing↗

Gene and enhancer trap tagging of vascular-expressed genes in poplar trees.

We report a gene discovery system for poplar trees based on gene and enhancer traps. Gene and enhancer trap vectors carrying the beta-glucuronidase (GUS) reporter gene were inserted into the poplar genome via Agrobacterium tumefaciens transformation, where they reveal the expression pattern of genes at or near the insertion sites. Because GUS expression phenotypes are dominant and are scored in primary transformants, this system does not require rounds of sexual recombination, a typical barrier to developmental genetic studies in trees. Gene and enhancer trap lines defining genes expressed during primary and secondary vascular development were identified and characterized. Collectively, the vascular gene expression patterns revealed that approximately 40% of genes expressed in leaves were expressed exclusively in the veins, indicating that a large set of genes is required for vascular development and function. Also, significant overlap was found between the sets of genes responsible for development and function of secondary vascular tissues of stems and primary vascular tissues in other organs of the plant, likely reflecting the common evolutionary origin of these tissues. Chromosomal DNA flanking insertion sites was amplified by thermal asymmetric interlaced PCR and sequenced and used to identify insertion sites by reference to the nascent Populus trichocarpa genome sequence. Extension of the system was demonstrated through isolation of full-length cDNAs for five genes of interest, including a new class of vascular-expressed gene tagged by enhancer trap line cET-1-pop1-145. Poplar gene and enhancer traps provide a new resource that allows plant biologists to directly reference the poplar genome sequence and identify novel genes of interest in forest biology.

Amino Acid Sequence↗