Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comprehensive genomic profiling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Molecular profiling of breast cancer.

PURPOSE OF REVIEW: This review is a comprehensive survey of molecular-profiling literature published since 2004. RECENT FINDINGS: More microarray-based gene-expression profiles that are prognostic for breast cancer have been published, strengthening the possibility that the microarray gene-expression profile may indeed provide clinically meaningful results. Requirement for snap-frozen tissue, however, will continue to be a limiting factor in clinical application. Results from a multicenter validation study were less spectacular than the original findings. A prognostic model based on classical markers performed well in a comparative study. Further clinical validation, with a large sample size, is needed. A prognostic gene-expression profile of 21 genes, which can be assayed using routinely processed formalin-fixed paraffin-embedded tumor tissue, has been introduced and this assay has also been shown to correlate with degree of benefit from chemotherapy. Two large clinical trials to validate gene-expression-based assays are to be launched in North America (TAILORx) and the European Union (MINDACCT). The usefulness of these genomic tools is still being debated, because clinicopathologic factors also are still important. SUMMARY: Gene-expression-based prognostic tests are now available as commercial reference laboratory tests. Their successful implementation will depend on the seamless integration with existing clinicopathologic markers.

Antineoplastic Agents↗

Comparative Analysis of Volatile Compounds, Amino Acids, Fatty Acids, and Lipidomic Profiles in Thigh Muscles of Commercial Arbor Acres (AA) Broilers and Indigenous Chengkou and Langshan Chickens.

Flavor-related compounds and nutritional components of chicken meat vary among different breeds, but comprehensive comparisons of these characteristics between commercial and indigenous chickens remain insufficiently characterized. In this study, three chicken breeds (Arbor Acres, Chengkou, and Langshan) were slaughtered at their respective market ages, and the volatile flavor compounds, amino acids, fatty acids, and lipidomic profiles of thigh muscle were analyzed to investigate breed-associated differences in flavor-related and nutritional characteristics. Langshan chickens exhibited the highest total volatile compound content and also had the highest total amino acid levels, with significantly higher contents of umami and sweet amino acids. In addition, both indigenous breeds showed higher levels of arachidonic acid (C20:4n6) than Arbor Acres broilers, while Chengkou chickens had the highest content of docosahexaenoic acid (DHA, C22:6n3). Lipidomic analysis identified 787 lipids, with glycerophospholipids and sphingolipids as the predominant classes. Differential lipid analysis revealed that Langshan chickens had 38 upregulated lipids compared with Arbor Acres chickens, while Chengkou chickens exhibited 258 differential lipids relative to Arbor Acres chickens. Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis indicated that these differential lipids were mainly associated with glycerolipid, sphingolipid, and glycerophospholipid metabolism. Correlation analysis further revealed significant associations between specific lipids and flavor-related compounds, amino acids, and fatty acids, suggesting their potential roles in breed-associated differences. Overall, this study demonstrates that indigenous chicken breeds possess distinct flavor-related and nutritional profiles compared with commercial Arbor Acres broilers and provides valuable insights into breed-associated differences in chicken meat characteristics.

amino acids↗

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks↗

Proteomic analysis of Acinetobacter lwoffii K24 by 2-D gel electrophoresis and electrospray ionization quadrupole-time of flight mass spectrometry.

The MS/MS analysis by Electrospray ionization quadrupole-time of flight mass spectrometry (ESI-Q-TOF MS) was applied to identify proteins in proteome analysis of bacteria whose genomes are not known. The protein identification by ESI-Q-TOF MS was performed sequentially by database search and then de novo sequencing using MS/MS spectra. Soil bacteria having unanalyzed genome, Acinetobacter lwoffii K24 is an aniline degrading bacterium. In this report, we present the results of a comparison between the proteome profile of A. lwoffii K24 cultured in aniline- or succinate-containing media. Protein analysis was performed using two-dimensional gel electrophoresis (2-DE) with pH 3-10 immobilized pH gradient (IPG) strips followed by ESI-Q-TOF MS. More than 780 protein spots were detected by 2-DE from the soluble proteome. Forty-eight of these proteins were expressed exclusively in aniline cultured bacteria, and 81 proteins increased and 162 proteins decreased in aniline-cultured versus succinate cultured A. lwoffii K24. Internal amino acid sequences of 43 major protein spots were successfully determined by ESI-Q-TOF MS to try to identify the bacterial proteins responding to aniline culture condition. Since the A. lwoffii K24 genome is not yet sequenced, many proteins were found to be hypothetical. Comparative proteome analysis of the insoluble protein fractions showed that one novel protein that was strongly induced by succinate-cultured A. lwoffii K24 was repressed under aniline culture conditions. These results suggest that comprehensive analysis of bacterial proteomes by 2-DE and amino acid sequence analysis by ESI-Q-TOF MS is useful for understanding induced novel proteins of biodegrading bacteria.

Acinetobacter↗

A comparative phylogenetic approach for dating whole genome duplication events.

MOTIVATION: Whole genome duplications have played a major role in determining the structure of eukaryotic genomes. Current evidence revealing large blocks of duplicated chromatin yields new insights into the evolutionary history of species, but also presents a major challenge for researchers attempting to utilize comparative genomics techniques. Understanding the timing of duplication events relative to divergence among taxa is critical to accurate and comprehensive cross-species comparisons. RESULTS: We describe a large-scale approach to estimate the timing of duplication events in a phylogenetic context. The methodology has been previously utilized for analysis of Arabidopsis and Saccharomyces duplication events. This new implementation provides a more flexible and reusable framework for these analyses. Scripts written in the Python programming language drive a number of freely available bioinformatics programs, creating a no-cost tool for researchers. The usefulness of the approach is demonstrated through genome-scale analysis of Arabidopsis and Oryza (rice) duplications. AVAILABILITY: Software and documentation are freely available from http://plantgenome.agtec.uga.edu/bioinformatics/dating/

Algorithms↗

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family↗

Genome-wide expression analysis in Corynebacterium glutamicum using DNA microarrays.

DNA microarray technology has become an important research tool for microbiology and biotechnology as it allows for comprehensive DNA and RNA analyses to characterize genetic diversity and gene expression in a genome-wide manner. DNA microarrays have been applied extensively to study the biology of many bacteria including Mycobacterium tuberculosis, but only recently have they been used for the related high-GC Gram-positive Corynebacterium glutamicum, which is widely used for biotechnological amino acid production. Besides the design and generation of microarrays as well as their use in hybridization experiments and subsequent data analysis, recent applications of DNA microarray technology in C. glutamicum including the characterization of ribose-specific gene expression and the valine stress response will be described. Emerging perspectives of functional genomics to enlarge our insight into fundamental biology of C. glutamicum and their impact on applied biotechnology will be discussed.

Corynebacterium↗

Visualization methods for statistical analysis of microarray clusters.

BACKGROUND: The most common method of identifying groups of functionally related genes in microarray data is to apply a clustering algorithm. However, it is impossible to determine which clustering algorithm is most appropriate to apply, and it is difficult to verify the results of any algorithm due to the lack of a gold-standard. Appropriate data visualization tools can aid this analysis process, but existing visualization methods do not specifically address this issue. RESULTS: We present several visualization techniques that incorporate meaningful statistics that are noise-robust for the purpose of analyzing the results of clustering algorithms on microarray data. This includes a rank-based visualization method that is more robust to noise, a difference display method to aid assessments of cluster quality and detection of outliers, and a projection of high dimensional data into a three dimensional space in order to examine relationships between clusters. Our methods are interactive and are dynamically linked together for comprehensive analysis. Further, our approach applies to both protein and gene expression microarrays, and our architecture is scalable for use on both desktop/laptop screens and large-scale display devices. This methodology is implemented in GeneVAnD (Genomic Visual ANalysis of Datasets) and is available at http://function.princeton.edu/GeneVAnD. CONCLUSION: Incorporating relevant statistical information into data visualizations is key for analysis of large biological datasets, particularly because of high levels of noise and the lack of a gold-standard for comparisons. We developed several new visualization techniques and demonstrated their effectiveness for evaluating cluster quality and relationships between clusters.

Algorithms↗

3D modelling of gene expression patterns.

The current genome-sequencing projects provide "word indices" of the book of life. A central post-genomic question will be how these words are three-dimensionally deployed in the generation of organism form. Gene expression studies of developing organisms contribute an increasing wealth of snapshot data on the activation of individual genes at selected locations and single moments in the developmental process. However, a comprehensive understanding of the dynamic activation of multiple genes and their functional role in controlling the 3D processes of collective cell behaviour, pattern formation and morphogenesis, requires special tools for a systematic description of spatio-temporal patterns of gene activation and the ensuing phenotypic effects. This article concentrates on new, computer-based tools for the 3D analysis of gene expression patterns in embryonic development and their use for the systematic establishment of comprehensive gene expression maps.

Animals↗

Fulfilling the promise: drug discovery in the post-genomic era.

The genomic era has brought with it a basic change in experimentation, enabling researchers to look more comprehensively at biological systems. The sequencing of the human genome coupled with advances in automation and parallelization technologies have afforded a fundamental transformation in the drug target discovery paradigm, towards systematic whole genome and proteome analyses. In conjunction with novel proteomic techniques, genome-wide annotation of function in cellular models is possible. Overlaying data derived from whole genome sequence, expression and functional analysis will facilitate the identification of causal genes in disease and significantly streamline the target validation process. Moreover, several parallel technological advances in small molecule screening have resulted in the development of expeditious and powerful platforms for elucidating inhibitors of protein or pathway function. Conversely, high-throughput and automated systems are currently being used to identify targets of orphan small molecules. The consolidation of these emerging functional genomics and drug discovery technologies promises to reap the fruits of the genomic revolution.

Animals↗

Pathologic diagnosis of thyroid nodules with preoperatively detected tumor protein 53 mutations: A tricenter series of 32 cases.

BACKGROUND: Mutations in the tumor protein 53 gene (TP53 mutation) in thyroid nodules are quoted to confer a high (80%) probability of malignancy when detected on the ThyroSeq v3 genomic classifier if associated with other molecular alterations. However, TP53 mutation also occurs in benign and low-risk thyroid neoplasms. Thus, the risk of malignancy in nodules harboring TP53 mutation is not well characterized. METHODS: Of 4,575 molecularly profiled preoperative fine-needle aspiration samples, 36 (0.8%) were identified harboring TP53 mutation. The study included 32 cases in which the pathology diagnosis was obtained from surgical specimens. RESULTS: The reviewed diagnosis was benign/low-risk neoplasms in 11 (34%), carcinoma-American Thyroid Association low risk of recurrence in 7 (22%), carcinoma-American Thyroid Association low-intermediate risk in 6 (19%), and carcinoma-American Thyroid Association high risk in 8 (25%). In the entire cohort and the indeterminate fine-needle aspiration category (Bethesda III-IV), the risk of malignancy was 66% and 56%, respectively. All 11 cases with a reviewed diagnosis of benign or low-risk neoplasms had their tumor capsule submitted entirely for histologic examination, and 64% had total thyroidectomy. In 30 cases comprehensively molecularly profiled, the molecular alterations were substratified into 4 groups: TP53 mutation alone (n = 5, 17%), TP53 mutation with copy number alteration (n = 9, 30%), TP53 mutation with other mutations but no copy number alteration (n = 9, 30%), and TP53 mutation with other mutations and copy number alteration (n = 7, 23%). The risk of malignancy for each group was 40%, 44%, 67%, and 100%, respectively. The frequency of American Thyroid Association-high-risk malignancy, which would often lead to a recommendation for total thyroidectomy, was 20%, 11%, 22%, and 43%, respectively. The risk of malignancy was higher in cases with additional mutations and copy number alteration (7/7, 100%) than in those with TP53 alone or with concomitant copy number alteration only (6/14, 43%) (P = .018). CONCLUSION: Thirty four percent of nodules with TP53 mutation with or without concomitant molecular alterations were benign/low-risk thyroid neoplasms and treated by total thyroidectomy in the majority of cases. Risk of malignancy increased significantly to 100% when TP53 mutation co-occurred with other mutations and copy number alterations. Since American Thyroid Association high-risk carcinomas were found in only 25% of TP53-mutated nodules, thyroid lobectomy may be considered as the initial treatment, in the appropriate clinical context.

Journal Article↗

Automatic pathway building in biological association networks.

BACKGROUND: Scientific literature is a source of the most reliable and comprehensive knowledge about molecular interaction networks. Formalization of this knowledge is necessary for computational analysis and is achieved by automatic fact extraction using various text-mining algorithms. Most of these techniques suffer from high false positive rates and redundancy of the extracted information. The extracted facts form a large network with no pathways defined. RESULTS: We describe the methodology for automatic curation of Biological Association Networks (BANs) derived by a natural language processing technology called Medscan. The curated data is used for automatic pathway reconstruction. The algorithm for the reconstruction of signaling pathways is also described and validated by comparison with manually curated pathways and tissue-specific gene expression profiles. CONCLUSION: Biological Association Networks extracted by MedScan technology contain sufficient information for constructing thousands of mammalian signaling pathways for multiple tissues. The automatically curated MedScan data is adequate for automatic generation of good quality signaling networks. The automatically generated Regulome pathways and manually curated pathways used for their validation are available free in the ResNetCore database from Ariadne Genomics, Inc. 1. The pathways can be viewed and analyzed through the use of a free demo version of PathwayStudio software. The Medscan technology is also available for evaluation using the free demo version of PathwayStudio software.

Databases, Bibliographic↗

Combined M-FISH and CGH analysis allows comprehensive description of genetic alterations in neuroblastoma cell lines.

Cancer cell lines are essential gene discovery tools and have often served as models in genetic and functional studies of particular tumor types. One of the future challenges is comparison and interpretation of gene expression data with the available knowledge on the genomic abnormalities in these cell lines. In this context, accurate description of these genomic abnormalities is required. Here, we show that a combination of M-FISH with banding analysis, standard FISH, and CGH allowed a detailed description of the genetic alterations in 16 neuroblastoma cell lines. In total, 14 cryptic chromosome rearrangements were detected, including a balanced t(2;4)(p24.3;q34.3) translocation in cell line NBL-S, with the 2p24 breakpoint located at about 40 kb from MYCN. The chromosomal origin of 22 marker chromosomes and 41 cytogenetically undefined translocated segments was determined. Chromosome arm 2 short arm translocations were observed in six cell lines (38%) with and five (31%) without MYCN amplification, leading to partial chromosome arm 2p gain in all but one cell line and loss of material in the various partner chromosomes, including 1p and 11q. These 2p gains were often masked in the GGH profiles due to MYCN amplification. The commonly overrepresented region was chromosome segment 2pter-2p22, which contains the MYCN gene, and five out of eleven 2p breakpoints clustered to the interface of chromosome bands 2p16 and 2p21. In neuroblastoma cell line SJNB-12, with double minutes (dmins) but no MYCN amplification, the dmins were shown to be derived from 16q22-q23 sequences. The ATBF1 gene, an AT-binding transcription factor involved in normal neurogenesis and located at 16q22.2, was shown to be present in the amplicon. This is the first report describing the possible implication of ATBF1 in neuroblastoma cells. We conclude that a combined approach of M-FISH, cytogenetics, and CGH allowed a more complete and accurate description of the genetic alterations occurring in the investigated cell lines.

Chromosome Painting↗

Cardiac transcriptional response to acute and chronic angiotensin II treatments.

Exposure of experimental animals to increased angiotensin II (ANG II) induces hypertension associated with cardiac hypertrophy, inflammation, and myocardial necrosis and fibrosis. Some of the most effective antihypertensive treatments are those that antagonize ANG II. We investigated cardiac gene expression in response to acute (24 h) and chronic (14 day) infusion of ANG II in mice; 24-h treatment induces hypertension, and 14-day treatment induces hypertension and extensive cardiac hypertrophy and necrosis. For genes differentially expressed in response to ANG II treatment, we tested for significant regulation of pathways, based on Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Microarray Pathway Profiler (GenMAPP) databases, as well as functional classes based on Gene Ontology (GO) terms. Both acute and chronic ANG II treatments resulted in decreased expression of mitochondrial metabolic genes, notably those for the electron transport chain and Krebs-TCA cycle; chronic ANG II treatment also resulted in decreased expression of genes involved in fatty acid metabolism. In contrast, genes involved in protein translation and ribosomal activity increased expression following both acute and chronic ANG II treatments. Some classes of genes showed differential response between acute and chronic ANG II treatments. Acute treatment increased expression of genes involved in oxidative stress and amino acid metabolism, whereas chronic treatments increased cytoskeletal and extracellular matrix genes, second messenger cascades responsive to ANG II, and amyloidosis genes. Although a functional linkage between Alzheimer disease, hypertension, and high cholesterol has been previously documented in studies of brain tissue, this is the first demonstration of induction of Alzheimer disease pathways by hypertension in heart tissue. This study provides the most comprehensive available survey of gene expression changes in response to acute and chronic ANG II treatment, verifying results from disparate studies, and suggests mechanisms that provide novel insight into the etiology of hypertensive heart disease and possible therapeutic interventions that may help to mitigate its effects.

Angiotensin II↗

Mouse cardiac surgery: comprehensive techniques for the generation of mouse models of human diseases and their application for genomic studies.

Mouse models mimicking human diseases are important tools in trying to understand the underlying mechanisms of many disease states. Several surgical models have been described that mimic human myocardial infarction (MI) and pressure-overload-induced cardiac hypertrophy. However, there are very few detailed descriptions for performing these surgical techniques in mice. Consequently, the number of laboratories that are proficient in performing cardiac surgical procedures in mice has been limited. Microarray technologies measure the expression of thousands of genes simultaneously, allowing for the identification of genes and pathways that may potentially be involved in the disease process. The statistical analysis of microarray experiments is highly influenced by the amount of variability in the experiment. To keep the number of required independent biological replicates and the associated costs of the study to a minimum, it is critical to minimize experimental variability by optimizing the surgical procedures. The aim of this publication was to provide a detailed description of techniques required to perform mouse cardiac surgery, such that these models can be utilized for genomic studies. A description of three major surgical procedures has been provided: 1) aortic constriction, 2) pulmonary artery banding, 3) MI (including ischemia-reperfusion). Emphasis has been placed on technical procedures with the inclusion of thorough descriptions of all equipment and devices employed in surgery, as well as the application of such techniques for expression profiling studies. The cardiac surgical techniques described have been, and will continue to be, important for elucidating the molecular mechanisms of cardiac hypertrophy and failure with high-throughput technology.

Anesthesia↗

In vivo expression profile of an endothelial nitric oxide synthase promoter-reporter transgene.

Endothelium-derived nitric oxide (NO) is primarily attributable to constitutive expression of the endothelial nitric oxide synthase (eNOS) gene. Although a more comprehensive understanding of transcriptional regulation of eNOS is emerging with respect to in vitro regulatory pathways, their relevance in vivo warrants assessment. In this regard, promoter-reporter insertional transgenic murine lines were created containing 5,200 bp of the native murine eNOS promoter directing transcription of nuclear-localized beta-galactosidase. Examination of beta-galactosidase expression in heart, lung, kidney, liver, spleen, and brain of adult mice demonstrated robust signal in large and medium-sized blood vessels. Small arterioles, capillaries, and venules of the microvasculature were notably negative, with the exception of the vasa recta of the medullary circulation of the kidney, which was strongly positive. Only in the brain was the reporter expressed in non-endothelial cell types, such as the CA1 region of the hippocampus. Epithelial cells of the bronchi, bronchioles, and alveoli were scored as negative, as was renal tubular epithelium. Cardiac myocytes, skeletal muscle, and smooth muscle of both vascular and nonvascular sources failed to demonstrate beta-galactosidase staining. Expression was uniform across multiple founders and was not significantly affected by genomic integration site. These transgenic eNOS promoter-reporter lines will be a valuable resource for ongoing studies addressing the regulated expression of eNOS in vivo in both health and disease.

Animals↗

Isolation of a library of target-sites for sequence specific DNA binding proteins from chick embryonic heart: a potential tool for identifying novel transcriptional regulators involved in embryonic development.

Enormity of the metazoan genomes and divergence in their regulation impose a serious constraint on the comprehensive understanding of context specific gene regulation. DNA elements located in the promoter, enhancer, and other regulatory regions of the genome dictate the temporal and spatial patterns of gene activities. However, owing to the diminutive and variable nature of the regulatory DNA elements, their identification and location remains a major challenge. We have developed an efficient strategy for isolating a repertoire of target sites for sequence specific DNA binding proteins from embryonic chick heart. A comprehensive library of such sequences was constructed and authenticated using various parameters including in silico determination of functional binding sites. This approach, therefore, for the first time, established an experimental and conceptual framework for defining the entire repertoire of functional DNA elements in any cellular context.

Amino Acid Sequence↗

Evolutionarily conserved regions and hydrophobic contacts at the superfamily level: The case of the fold-type I, pyridoxal-5'-phosphate-dependent enzymes.

The wealth of biological information provided by structural and genomic projects opens new prospects of understanding life and evolution at the molecular level. In this work, it is shown how computational approaches can be exploited to pinpoint protein structural features that remain invariant upon long evolutionary periods in the fold-type I, PLP-dependent enzymes. A nonredundant set of 23 superposed crystallographic structures belonging to this superfamily was built. Members of this family typically display high-structural conservation despite low-sequence identity. For each structure, a multiple-sequence alignment of orthologous sequences was obtained, and the 23 alignments were merged using the structural information to obtain a comprehensive multiple alignment of 921 sequences of fold-type I enzymes. The structurally conserved regions (SCRs), the evolutionarily conserved residues, and the conserved hydrophobic contacts (CHCs) were extracted from this data set, using both sequence and structural information. The results of this study identified a structural pattern of hydrophobic contacts shared by all of the superfamily members of fold-type I enzymes and involved in native interactions. This profile highlights the presence of a nucleus for this fold, in which residues participating in the most conserved native interactions exhibit preferential evolutionary conservation, that correlates significantly (r = 0.70) with the extent of mean hydrophobic contact value of their apolar fraction.

Conserved Sequence↗