Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

The prostate expression database (PEDB): status and enhancements in 2000.

The Prostate Expression Database (PEDB) is an online resource designed to access and analyze gene expression information derived from the human prostate. PEDB archives >55 000 expressed sequence tags (ESTs) from 43 cDNA libraries in a curated relational database that provides detailed library information including tissue source, library construction methods, sequence diversity and sequence abundance. The differential expression of each EST species can be viewed across all libraries using a Virtual Expression Analysis Tool (VEAT), a graphical user interface written in Java for intra- and inter-library species comparisons. Recent enhancements to PEDB include: (i) the functional categorization of annotated EST assemblies using a classification scheme developed at The Institute for Genome Research; (ii) catalogs of expressed genes in specific prostate tissue sources designated as transcriptomes; and (iii) the addition of prostate proteome information derived from two-dimensional electrophoreses and mass spectrometry of prostate cancer cell lines. PEDB may be accessed via the WWW at http://www.mbt.washington.edu/PEDB/

Databases, Factual↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Patterns of Genomic Divergence and Introgression in Two Primulina Hybrid Zones.

Hybrid zones have long been promoted as natural laboratories for understanding the mechanisms of speciation. Multiple or replicated hybrid zones are particularly informative, as they allow for assessing the consistency of genomic divergence and introgression across different environmental contexts and demographic histories, thereby improving our understanding of the factors that drive or hinder speciation on a broader scale. Here, using whole-genome resequencing data, we compare the patterns of genomic divergence and introgression in two Primulina hybrid zones. We found that genomic divergence in both hybrid zones is largely shaped by neutral processes, with only a few genomic regions showing signatures of balancing or lineage-specific selection. Genomic cline analyses identified numerous SNPs that showed significantly steeper clines and biased centres than the genome-wide expectation in both hybrid zones, consistent with the existence of reproductive barriers. Within regions of restricted gene flow, we identified 21 genes shared between the two hybrid zones. Annotation of gene function revealed that several genes are involved in reproductive processes. In addition, many zone-specific outlier loci were linked to genes associated with pollen and flower development, suggesting that these barriers may contribute to reproductive isolation under localised ecological conditions. Overall, these findings suggest that while certain reproductive barriers remain consistent across independent hybrid zones, others may be contingent on local environmental contexts. Our results demonstrate that both general and zone-specific mechanisms contribute to reproductive isolation in Primulina, providing empirical evidence that some genomic barriers recur across independent hybrid zones while others arise through localised adaptation.

Lamiales↗

Hypertrophic cardiomyopathy: a genome-wide association meta-analysis and polygenic risk score.

BACKGROUND: Hypertrophic cardiomyopathy (HCM) is a heritable trait with marked variability in expression and outcomes. Our aims were to discover new genetic loci associated with HCM and to test the effect of a new polygenic risk score (PRS) on incidence, phenotype and outcomes stratified by genotype status. METHODS: A discovery genome-wide association study (GWAS) was performed on 2284 HCM cases and 4525 controls. Two fixed-effects meta-analyses combined our discovery GWAS with single-trait and multi-trait results from a published study. Discovered loci underwent comprehensive bioinformatic analysis including functional and druggability annotations. A PRS using loci from the two meta-analyses was evaluated for association with HCM diagnosis in 411 213 individuals from UK Biobank (UKBB); imaging phenotypes in individuals without HCM; a composite endpoint (including all-cause mortality and transplantation); and sudden cardiac death (SCD) in 1756 HCM cases. PRS analyses were stratified by genotype status. RESULTS: Three loci were found in the discovery GWAS (BAG3, FHOD3 and novel locus PPP1R3A). In the meta-analyses, 70 unique loci were identified, four novel (MYPN, YWHAE, NOS1AP and OBSCN). Bioinformatic analyses identified NOS1AP as a candidate HCM gene. A new PRS was significantly associated with HCM diagnosis (HR=3.19, 95% CI 2.46 to 4.14 for top 5% vs lower 95%; HR=1.88, 95% CI 1.72 to 2.06 per SD increase). Significant associations were found between PRS and greater left ventricular (LV) wall thickness and higher LV ejection fraction in UKBB participants without HCM. Genotype-negative HCM cases in the top 20% of the PRS distribution had an increased risk of SCD (HR=2.72, 95% CI 1.03 to 7.17). CONCLUSIONS: We report novel HCM loci. A new PRS predicted the risk of HCM development and associated imaging characteristics in the UKBB and outcomes in an HCM cohort.

Cardiomyopathies↗

Genetic insights into the relationship between age at menarche and mental health-related phenotypes.

BACKGROUND: Multiple observational studies have reported associations between age at menarche (AAM) and mental health problems, yet their shared genetic architecture remains poorly characterized. METHODS: We leveraged genome-wide association study summary statistics for AAM and 15 mental health-related phenotypes. We conducted a multi-method integrative analysis encompassing linkage disequilibrium score regression, pleiotropic analysis under the composite null hypothesis, functional mapping and annotation, multi-marker analysis of genomic annotation, pathway enrichment, and bidirectional two-sample Mendelian randomization (MR) to explore shared genetic architecture and potential causal relationships. RESULTS: Our study identified significant genetic correlations between AAM and eight mental health-related phenotypes (miserableness, fed-up feelings, nervous feelings, ever thought that life is not worth living, ever self-harmed, depression, ever smoker, and age started smoking in former smokers). A total of 155 pleiotropic loci, 18 colocalized loci (e.g., 6q16.3), and 203 pleiotropic genes (e.g., LIN28B) were identified. These genes are expressed in multiple regions, including the cerebral cortex and hypothalamus, and are involved in various biological processes and signaling pathways. Additionally, MR analysis revealed causal associations between AAM and 5 mental health-related phenotypes (mood swings, miserableness, fed-up feelings, and age at which smokers started smoking in former/current smokers). CONCLUSIONS: Our study revealed extensive genetic associations between AAM and mental health-related phenotypes, and further explored the potential causal relationships between them. These findings enhance our understanding of the relationship from a genetic perspective and establish a foundation for future research to explore the biological pathways and environmental interactions contributing to these associations.

Genome-Wide Association Study↗

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning↗

Development and validation of whole-genome SSR markers in sugar beet (Beta vulgaris L.).

Sugar beet (Beta vulgaris L.) is an important sugar and cash crop worldwide. To systematically characterize SSR (Simple Sequence Repeat) loci across sugar beet chromosomes and enable the precise identification of germplasm resources, this study conducted a genome-wide scan for SSR loci, analyzed their distribution patterns, and determined their genotypes using resequencing data from 123 sugar beet varieties. The results revealed an abundance of SSR loci in the sugar beet genome, with a total of 135, 379 identified, from which 135, 344 pairs of SSR primers were designed (135, 344 primer pairs successfully designed; 35 loci failed to meet design criteria). Specifically, 31, 748 primer pairs were designed based on SSRs located in unassigned scaffolds, and 103, 596 primer pairs from SSRs assigned to the nine chromosomes. Through bioinformatic analysis, we identified 28, 768 SSR primers located in multi-copy genes with PIC (Polymorphism Information Content) ≥ 0.5, and 2, 326 SSR markers located in single-copy genes residing in various genic regions (among which 543 had PIC ≥ 0.5, with the highest reaching 0.776). PCR (Polymerase Chain Reaction) validation confirmed 20 robust and polymorphic markers producing clear and reproducible bands. Among them, 10 SSR primers located in multi-copy genes exhibited three or more polymorphic types, and 10 markers located in single-copy genes displayed 2-3 polymorphic types. The most polymorphic marker, YCD-4-2, detected 11 polymorphic types across 48 varieties. Furthermore, to explore markers with potential functional significance, we annotated the genes harboring SSR markers located in single-copy genes. The results showed that 1, 264 SSRs located in single-copy genes were localized to 967 genes, which are significantly enriched in pathways related to carbohydrate metabolism, stress responses, and plant-pathogen interactions. The 20 validated markers and the 2, 326 SSRs located in single-copy genes provided in this study can be directly applied to fingerprinting of sugar beet varieties, seed purity testing, and marker-assisted selection, thus representing a practical resource for molecular breeding.

genome-wide↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.

Measuring in a quantitative, statistical sense the degree to which structural and functional information can be "transferred" between pairs of related protein sequences at various levels of similarity is an essential prerequisite for robust genome annotation. To this end, we performed pairwise sequence, structure and function comparisons on approximately 30,000 pairs of protein domains with known structure and function. Our domain pairs, which are constructed according to the SCOP fold classification, range in similarity from just sharing a fold, to being nearly identical. Our results show that traditional scores for sequence and structure similarity have the same basic exponential relationship as observed previously, with structural divergence, measured in RMS, being exponentially related to sequence divergence, measured in percent identity. However, as the scale of our survey is much larger than any previous investigations, our results have greater statistical weight and precision. We have been able to express the relationship of sequence and structure similarity using more "modern scores," such as Smith-Waterman alignment scores and probabilistic P-values for both sequence and structure comparison. These modern scores address some of the problems with traditional scores, such as determining a conserved core and correcting for length dependency; they enable us to phrase the sequence-structure relationship in more precise and accurate terms. We found that the basic exponential sequence-structure relationship is very general: the same essential relationship is found in the different secondary-structure classes and is evident in all the scoring schemes. To relate function to sequence and structure we assigned various levels of functional similarity to the domain pairs, based on a simple functional classification scheme. This scheme was constructed by combining and augmenting annotations in the enzyme and fly functional classifications and comparing subsets of these to the Escherichia coli and yeast classifications. We found sigmoidal relationships between similarity in function and sequence, with clear thresholds for different levels of functional conservation. For pairs of domains that share the same fold, precise function appears to be conserved down to approximately 40 % sequence identity, whereas broad functional class is conserved to approximately 25 %. Interestingly, percent identity is more effective at quantifying functional conservation than the more modern scores (e.g. P-values). Results of all the pairwise comparisons and our combined functional classification scheme for protein structures can be accessed from a web database at http://bioinfo.mbb.yale.edu/alignCopyright 2000 Academic Press.

Animals↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Assigning function to CDS through qualified query answering: beyond alignment and motifs.

In this paper, we show how to use qualitative query answering to annotate CDS-to-function relationships with confidence in the score, confidence in the tool, and confidence in the decision about the function. The system, implemented in Prolog, provides users with a powerful tool to analyze large quantities of data that have been produce by multiple sequence analysis programs. Using qualified query answering techniques, users can easily change the criteria for how tools reinforce each other and for how numbers of occurrences of particular functions reinforce each other. They can also alter how different scores for different tools are categorized.

Animals↗

In silico predictions of Escherichia coli metabolic capabilities are consistent with experimental data.

A significant goal in the post-genome era is to relate the annotated genome sequence to the physiological functions of a cell. Working from the annotated genome sequence, as well as biochemical and physiological information, it is possible to reconstruct complete metabolic networks. Furthermore, computational methods have been developed to interpret and predict the optimal performance of a metabolic network under a range of growth conditions. We have tested the hypothesis that Escherichia coli uses its metabolism to grow at a maximal rate using the E. coli MG1655 metabolic reconstruction. Based on this hypothesis, we formulated experiments that describe the quantitative relationship between a primary carbon source (acetate or succinate) uptake rate, oxygen uptake rate, and maximal cellular growth rate. We found that the experimental data were consistent with the stated hypothesis, namely that the E. coli metabolic network is optimized to maximize growth under the experimental conditions considered. This study thus demonstrates how the combination of in silico and experimental biology can be used to obtain a quantitative genotype-phenotype relationship for metabolism in bacterial cells.

Acetates↗

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors↗

The TIGR gene indices: reconstruction and representation of expressed gene sequences.

Expressed sequence tags (ESTs) have provided a first glimpse of the collection of transcribed sequences in a variety of organisms. However, a careful analysis of this sequence data can provide significant additional functional, structural and evolutionary information. Our analysis of the public EST sequences, available through the TIGR Gene Indices (TGI; http://www.tigr.org/tdb/tdb.html ), is an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed for selected organisms by first clustering, then assembling EST and annotated gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, and to provide links between orthologous and paralogous genes.

Base Sequence↗