Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

Hypertrophic cardiomyopathy: a genome-wide association meta-analysis and polygenic risk score.

BACKGROUND: Hypertrophic cardiomyopathy (HCM) is a heritable trait with marked variability in expression and outcomes. Our aims were to discover new genetic loci associated with HCM and to test the effect of a new polygenic risk score (PRS) on incidence, phenotype and outcomes stratified by genotype status. METHODS: A discovery genome-wide association study (GWAS) was performed on 2284 HCM cases and 4525 controls. Two fixed-effects meta-analyses combined our discovery GWAS with single-trait and multi-trait results from a published study. Discovered loci underwent comprehensive bioinformatic analysis including functional and druggability annotations. A PRS using loci from the two meta-analyses was evaluated for association with HCM diagnosis in 411 213 individuals from UK Biobank (UKBB); imaging phenotypes in individuals without HCM; a composite endpoint (including all-cause mortality and transplantation); and sudden cardiac death (SCD) in 1756 HCM cases. PRS analyses were stratified by genotype status. RESULTS: Three loci were found in the discovery GWAS (BAG3, FHOD3 and novel locus PPP1R3A). In the meta-analyses, 70 unique loci were identified, four novel (MYPN, YWHAE, NOS1AP and OBSCN). Bioinformatic analyses identified NOS1AP as a candidate HCM gene. A new PRS was significantly associated with HCM diagnosis (HR=3.19, 95% CI 2.46 to 4.14 for top 5% vs lower 95%; HR=1.88, 95% CI 1.72 to 2.06 per SD increase). Significant associations were found between PRS and greater left ventricular (LV) wall thickness and higher LV ejection fraction in UKBB participants without HCM. Genotype-negative HCM cases in the top 20% of the PRS distribution had an increased risk of SCD (HR=2.72, 95% CI 1.03 to 7.17). CONCLUSIONS: We report novel HCM loci. A new PRS predicted the risk of HCM development and associated imaging characteristics in the UKBB and outcomes in an HCM cohort.

Cardiomyopathies↗

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

Deciphering cellular states of innate tumor drug responses.

BACKGROUND: The molecular mechanisms underlying innate tumor drug resistance, a major obstacle to successful cancer therapy, remain poorly understood. In colorectal cancer (CRC), molecular studies have focused on drug-selected tumor cell lines or individual candidate genes using samples derived from patients already treated with drugs, so that very little data are available prior to drug treatment. RESULTS: Transcriptional profiles of clinical samples collected from CRC patients prior to their exposure to a combined chemotherapy of folinic acid, 5-fluorouracil and irinotecan were established using microarrays. Vigilant experimental design, power simulations and robust statistics were used to restrain the rates of false negative and false positive hybridizations, allowing successful discrimination between drug resistance and sensitivity states with restricted sampling. A list of 679 genes was established that intrinsically differentiates, for the first time prior to drug exposure, subsequently diagnosed chemo-sensitive and resistant patients. Independent biological validation performed through quantitative PCR confirmed the expression pattern on two additional patients. Careful annotation of interconnected functional networks provided a unique representation of the cellular states underlying drug responses. CONCLUSION: Molecular interaction networks are described that provide a solid foundation on which to anchor working hypotheses about mechanisms underlying in vivo innate tumor drug responses. These broad-spectrum cellular signatures represent a starting point from which by-pass chemotherapy schemes, targeting simultaneously several of the molecular mechanisms involved, may be developed for critical therapeutic intervention in CRC patients. The demonstrated power of this research strategy makes it generally applicable to other physiological and pathological situations.

Antineoplastic Agents↗

Genetic insights into the relationship between age at menarche and mental health-related phenotypes.

BACKGROUND: Multiple observational studies have reported associations between age at menarche (AAM) and mental health problems, yet their shared genetic architecture remains poorly characterized. METHODS: We leveraged genome-wide association study summary statistics for AAM and 15 mental health-related phenotypes. We conducted a multi-method integrative analysis encompassing linkage disequilibrium score regression, pleiotropic analysis under the composite null hypothesis, functional mapping and annotation, multi-marker analysis of genomic annotation, pathway enrichment, and bidirectional two-sample Mendelian randomization (MR) to explore shared genetic architecture and potential causal relationships. RESULTS: Our study identified significant genetic correlations between AAM and eight mental health-related phenotypes (miserableness, fed-up feelings, nervous feelings, ever thought that life is not worth living, ever self-harmed, depression, ever smoker, and age started smoking in former smokers). A total of 155 pleiotropic loci, 18 colocalized loci (e.g., 6q16.3), and 203 pleiotropic genes (e.g., LIN28B) were identified. These genes are expressed in multiple regions, including the cerebral cortex and hypothalamus, and are involved in various biological processes and signaling pathways. Additionally, MR analysis revealed causal associations between AAM and 5 mental health-related phenotypes (mood swings, miserableness, fed-up feelings, and age at which smokers started smoking in former/current smokers). CONCLUSIONS: Our study revealed extensive genetic associations between AAM and mental health-related phenotypes, and further explored the potential causal relationships between them. These findings enhance our understanding of the relationship from a genetic perspective and establish a foundation for future research to explore the biological pathways and environmental interactions contributing to these associations.

Genome-Wide Association Study↗

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning↗

Development and validation of whole-genome SSR markers in sugar beet (Beta vulgaris L.).

Sugar beet (Beta vulgaris L.) is an important sugar and cash crop worldwide. To systematically characterize SSR (Simple Sequence Repeat) loci across sugar beet chromosomes and enable the precise identification of germplasm resources, this study conducted a genome-wide scan for SSR loci, analyzed their distribution patterns, and determined their genotypes using resequencing data from 123 sugar beet varieties. The results revealed an abundance of SSR loci in the sugar beet genome, with a total of 135, 379 identified, from which 135, 344 pairs of SSR primers were designed (135, 344 primer pairs successfully designed; 35 loci failed to meet design criteria). Specifically, 31, 748 primer pairs were designed based on SSRs located in unassigned scaffolds, and 103, 596 primer pairs from SSRs assigned to the nine chromosomes. Through bioinformatic analysis, we identified 28, 768 SSR primers located in multi-copy genes with PIC (Polymorphism Information Content) ≥ 0.5, and 2, 326 SSR markers located in single-copy genes residing in various genic regions (among which 543 had PIC ≥ 0.5, with the highest reaching 0.776). PCR (Polymerase Chain Reaction) validation confirmed 20 robust and polymorphic markers producing clear and reproducible bands. Among them, 10 SSR primers located in multi-copy genes exhibited three or more polymorphic types, and 10 markers located in single-copy genes displayed 2-3 polymorphic types. The most polymorphic marker, YCD-4-2, detected 11 polymorphic types across 48 varieties. Furthermore, to explore markers with potential functional significance, we annotated the genes harboring SSR markers located in single-copy genes. The results showed that 1, 264 SSRs located in single-copy genes were localized to 967 genes, which are significantly enriched in pathways related to carbohydrate metabolism, stress responses, and plant-pathogen interactions. The 20 validated markers and the 2, 326 SSRs located in single-copy genes provided in this study can be directly applied to fingerprinting of sugar beet varieties, seed purity testing, and marker-assisted selection, thus representing a practical resource for molecular breeding.

genome-wide↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗

The cell membrane and the struggle for life of lactic acid bacteria.

The major life-threatening event for lactic acid bacteria (LAB) in their natural environment is the depletion of their energy sources and LAB can survive such conditions only for a short period of time. During periods of starvation LAB can exploit optimally the potential energy sources in their environment usually by applying proton motive force generating membrane transport systems. These systems include in addition to the proton translocating F0F1-ATPase: a respiratory chain when hemin is present in the medium, electrogenic solute uptake and excretion systems, electrogenic lactate/proton symport and precursor/product exchange systems. Most of these metabolic energy-generating systems offer as additional bonus the prevention of a lethal decrease of the internal and external pH. LAB have limited biosynthetic capacities and rely heavily on the presence of essential components such as sources of amino acids in their environment. The uptake of amino acids requires a major fraction of the available metabolic energy of LAB. The metabolic energy cost of amino acid uptake can be reduced drastically by accumulating oligopeptides instead of the individual amino acids and by proton motive force-generating efflux of excessively accumulated amino acids. Other life-threatening conditions that LAB encounter in their environment are rapid changes in the osmolality and the exposure to cytotoxic compounds, including antibiotics. LAB respond to osmotic upshock or downshock by accumulating or releasing rapidly osmolytes such as glycine-betaine. The life-threatening presence of cytotoxic compounds, including antibiotics, is effectively counteracted by powerful drug extruding multidrug resistance systems. The number and variety of defense mechanisms in LAB is surprisingly high. Most defense mechanisms operate in the cytoplasmic membrane to control the internal environment and the energetic status of LAB. Annotation of the functions of the genes in the genomes of LAB will undoubtedly reveal additional defense mechanisms.

Biological Transport↗

YMD: a microarray database for large-scale gene expression analysis.

The use of microarray technology to perform parallel analysis of the expression pattern of a large number of genes in a single experiment has created a new frontier of medical research. The vast amount of gene expression data generated from multiple microarray experiments requires a robust database system that allows efficient data storage, retrieval, secure access, data dissemination, and integrated data analyses. To address the growing needs of microarray researchers at Yale and their collaborators, we have built the Yale Microarray Database (YMD). YMD is Web-accessible with the following features: (i) a Web program that tracks DNA samples between source plates and arrays, (ii) the capability of finding common genes/clones across different array platforms, (iii) an image file server, (iv) laboratory-based user management and access privileges, (v) project management, (vi) template data entry, (vii) linking gene expression data to annotation databases for functional analysis. YMD is currently being used on a pilot basis by several laboratories for different organisms and array platforms.

Databases, Nucleic Acid↗

TIGRFAMs and Genome Properties: tools for the assignment of molecular function and biological process in prokaryotic genomes.

TIGRFAMs is a collection of protein family definitions built to aid in high-throughput annotation of specific protein functions. Each family is based on a hidden Markov model (HMM), where both cutoff scores and membership in the seed alignment are chosen so that the HMMs can classify numerous proteins according to their specific molecular functions. Most TIGRFAMs models describe 'equivalog' families, where both orthology and lateral gene transfer may be part of the evolutionary history, but where a single molecular function has been conserved. The Genome Properties system contains a queriable set of metabolic reconstructions, genome metrics and extractions of information from the scientific literature. Its genome-by-genome assertions of whether or not specific structures, pathways or systems are present provide high-level conceptual descriptions of genomic content. These assertions enable comparative genomics, provide a meaningful biological context to aid in manual annotation, support assignments of Gene Ontology (GO) biological process terms and help validate HMM-based predictions of protein function. The Genome Properties system is particularly useful as a generator of phylogenetic profiles, through which new protein family functions may be discovered. The TIGRFAMs and Genome Properties systems can be accessed at http://www.tigr.org/TIGRFAMs and http://www.tigr.org/Genome_Properties.

Archaeal Proteins↗

The Drosophila melanogaster genome.

Drosophila's importance as a model organism made it an obvious choice to be among the first genomes sequenced, and the Release 1 sequence of the euchromatic portion of the genome was published in March 2000. This accomplishment demonstrated that a whole genome shotgun (WGS) strategy could produce a reliable metazoan genome sequence. Despite the attention to sequencing methods, the nucleotide sequence is just the starting point for genome-wide analyses; at a minimum, the genome sequence must be interpreted using expressed sequence tag (EST) and complementary DNA (cDNA) evidence and computational tools to identify genes and predict the structures of their RNA and protein products. The functions of these products and the manner in which their expression and activities are controlled must then be assessed-a much more challenging task with no clear endpoint that requires a wide variety of experimental and computational methods. We first review the current state of the Drosophila melanogaster genome sequence and its structural annotation and then briefly summarize some promising approaches that are being taken to achieve an initial functional annotation.

Animals↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

dictyBase, the model organism database for Dictyostelium discoideum.

dictyBase (http://dictybase.org) is the model organism database (MOD) for the social amoeba Dictyostelium discoideum. The unique biology and phylogenetic position of Dictyostelium offer a great opportunity to gain knowledge of processes not characterized in other organisms. The recent completion of the 34 MB genome sequence, together with the sizable scientific literature using Dictyostelium as a research organism, provided the necessary tools to create a well-annotated genome. dictyBase has leveraged software developed by the Saccharomyces Genome Database and the Generic Model Organism Database project. This has reduced the time required to develop a full-featured MOD and greatly facilitated our ability to focus on annotation and providing new functionality. We hope that manual curation of the Dictyostelium genome will facilitate the annotation of other genomes.

Animals↗

Quantifying structure-function uncertainty: a graph theoretical exploration into the origins and limitations of protein annotation.

Since the advent of investigations into structural genomics, research has focused on correctly identifying domain boundaries, as well as domain similarities and differences in the context of their evolutionary relationships. As the science of structural genomics ramps up adding more and more information into the databanks, questions about the accuracy and completeness of our classification and annotation systems appear on the forefront of this research. A central question of paramount importance is how structural similarity relates to functional similarity. Here, we begin to rigorously and quantitatively answer these questions by first exploring the consensus between the most common protein domain structure annotation databases CATH, SCOP and FSSP. Each of these databases explores the evolutionary relationships between protein domains using a combination of automatic and manual, structural and functional, continuous and discrete similarity measures. In order to examine the issue of consensus thoroughly, we build a generalized graph out of each of these databases and hierarchically cluster these graphs at interval thresholds. We then employ a distance measure to find regions of greatest overlap. Using this procedure we were able not only to enumerate the level of consensus between the different annotation systems, but also to define the graph-theoretical origins behind the annotation schema of class, family and superfamily by observing that the same thresholds that define the best consensus regions between FSSP, SCOP and CATH correspond to distinct, non-random phase-transitions in the structure comparison graph itself. To investigate the correspondence in divergence between structure and function further, we introduce a measure of functional entropy that calculates divergence in function space. First, we use this measure to calculate the general correlation between structural homology and functional proximity. We extend this analysis further by quantitatively calculating the average amount of functional information gained from our understanding of structural distance and the corollary inherent uncertainty that represents the theoretical limit of our ability to infer function from structural similarity. Finally we show how our measure of functional "entropy" translates into a more intuitive concept of functional annotation into similarity EC classes.

Biochemical Phenomena↗

Combining text mining and sequence analysis to discover protein functional regions.

Recently presented protein sequence classification models can identify relevant regions of the sequence. This observation has many potential applications to detecting functional regions of proteins. However, identifying such sequence regions automatically is difficult in practice, as relatively few types of information have enough annotated sequences to perform this analysis. Our approach addresses this data scarcity problem by combining text and sequence analysis. First, we train a text classifier over the explicit textual annotations available for some of the sequences in the dataset, and use the trained classifier to predict the class for the rest of the unlabeled sequences. We then train a joint sequence text classifier over the text contained in the functional annotations of the sequences, and the actual sequences in this larger, automatically extended dataset. Finally, we project the classifier onto the original sequences to determine the relevant regions of the sequences. We demonstrate the effectiveness of our approach by predicting protein sub-cellular localization and determining localization specific functional regions of these proteins.

Algorithms↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

In silico gene function prediction using ontology-based pattern identification.

MOTIVATION: With the emergence of genome-wide expression profiling data sets, the guilt by association (GBA) principle has been a cornerstone for deriving gene functional interpretations in silico. Given the limited success of traditional methods for producing clusters of genes with great amounts of functional similarity, new data-mining algorithms are required to fully exploit the potential of high-throughput genomic approaches. RESULTS: Ontology-based pattern identification (OPI) is a novel data-mining algorithm that systematically identifies expression patterns that best represent existing knowledge of gene function. Instead of relying on a universal threshold of expression similarity to define functionally related groups of genes, OPI finds the optimal analysis settings that yield gene expression patterns and gene lists that best predict gene function using the principle of GBA. We applied OPI to a publicly available gene expression data set on the life cycle of the malarial parasite Plasmodium falciparum and systematically annotated genes for 320 functional categories based on current Gene Ontology annotations. An ontology-based hierarchical tree of the 320 categories provided a systems-wide biological view of this important malarial parasite.

Algorithms↗

Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.

Measuring in a quantitative, statistical sense the degree to which structural and functional information can be "transferred" between pairs of related protein sequences at various levels of similarity is an essential prerequisite for robust genome annotation. To this end, we performed pairwise sequence, structure and function comparisons on approximately 30,000 pairs of protein domains with known structure and function. Our domain pairs, which are constructed according to the SCOP fold classification, range in similarity from just sharing a fold, to being nearly identical. Our results show that traditional scores for sequence and structure similarity have the same basic exponential relationship as observed previously, with structural divergence, measured in RMS, being exponentially related to sequence divergence, measured in percent identity. However, as the scale of our survey is much larger than any previous investigations, our results have greater statistical weight and precision. We have been able to express the relationship of sequence and structure similarity using more "modern scores," such as Smith-Waterman alignment scores and probabilistic P-values for both sequence and structure comparison. These modern scores address some of the problems with traditional scores, such as determining a conserved core and correcting for length dependency; they enable us to phrase the sequence-structure relationship in more precise and accurate terms. We found that the basic exponential sequence-structure relationship is very general: the same essential relationship is found in the different secondary-structure classes and is evident in all the scoring schemes. To relate function to sequence and structure we assigned various levels of functional similarity to the domain pairs, based on a simple functional classification scheme. This scheme was constructed by combining and augmenting annotations in the enzyme and fly functional classifications and comparing subsets of these to the Escherichia coli and yeast classifications. We found sigmoidal relationships between similarity in function and sequence, with clear thresholds for different levels of functional conservation. For pairs of domains that share the same fold, precise function appears to be conserved down to approximately 40 % sequence identity, whereas broad functional class is conserved to approximately 25 %. Interestingly, percent identity is more effective at quantifying functional conservation than the more modern scores (e.g. P-values). Results of all the pairwise comparisons and our combined functional classification scheme for protein structures can be accessed from a web database at http://bioinfo.mbb.yale.edu/alignCopyright 2000 Academic Press.

Animals↗