Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Classification and knowledge discovery in protein databases.

We consider the problem of classification in noisy, high-dimensional, and class-imbalanced protein datasets. In order to design a complete classification system, we use a three-stage machine learning framework consisting of a feature selection stage, a method addressing noise and class-imbalance, and a method for combining biologically related tasks through a prior-knowledge based clustering. In the first stage, we employ Fisher's permutation test as a feature selection filter. Comparisons with the alternative criteria show that it may be favorable for typical protein datasets. In the second stage, noise and class imbalance are addressed by using minority class over-sampling, majority class under-sampling, and ensemble learning. The performance of logistic regression models, decision trees, and neural networks is systematically evaluated. The experimental results show that in many cases ensembles of logistic regression classifiers may outperform more expressive models due to their robustness to noise and low sample density in a high-dimensional feature space. However, ensembles of neural networks may be the best solution for large datasets. In the third stage, we use prior knowledge to partition unlabeled data such that the class distributions among non-overlapping clusters significantly differ. In our experiments, training classifiers specialized to the class distributions of each cluster resulted in a further decrease in classification error.

Algorithms↗

Improving the classification of multiple disorders with problem decomposition.

Differential diagnosis of multiple disorders is a challenging problem in clinical medicine. According to the divide-and-conquer principle, this problem can be handled more effectively through decomposing it into a number of simpler sub-problems, each solved separately. We demonstrate the advantages of this approach using abductive network classifiers on the 6-class standard dermatology dataset. Three problem decomposition scenarios are investigated, including class decomposition and two hierarchical approaches based on clinical practice and class separability properties. Two-stage classification schemes based on hierarchical decomposition boost the classification accuracy from 91% for the single-classifier monolithic approach to 99%, matching the theoretical upper limit reported in the literature for the accuracy of classifying the dataset. Such models are also simpler, achieving up to 47% reduction in the number of input variables required, thus reducing the cost and improving the convenience of performing the medical diagnostic tests required. Automatic selection of only relevant inputs by the simpler abductive network models synthesized provides greater insight into the diagnosis problem and the diagnostic value of various disease markers. The problem decomposition approach helps plan more efficient diagnostic tests and provides improved support for the decision-making process. Findings are compared with established guidelines of clinical practice, results of data analysis, and outcomes of previous informatics-based studies on the dataset.

Algorithms↗

Patterns of intra-cluster correlation from primary care research to inform study design and analysis.

OBJECTIVE: To provide information concerning the magnitude of the intraclass correlation coefficient (ICC) for cluster-based studies set in primary care. STUDY DESIGN AND SETTING: Reanalysis of data from 31 cluster-based studies in primary care to estimate intraclass correlation coefficients from random effects models using maximum likelihood estimation. RESULTS: ICCs were estimated for 1,039 variables. The median ICC was 0.010 (interquartile range [IQR] 0 to 0.032, range 0 to 0.840). After adjusting for individual- and cluster-level characteristics, the median ICC was 0.005 (IQR 0 to 0.021). A given measure showed widely varying ICC estimates in different datasets. In six datasets, the ICCs for SF-36 physical functioning scale ranged from 0.001 to 0.055 and for SF-36 general health from 0 to 0.072. In four datasets, the ICC for systolic blood pressure ranged from 0 to 0.052 and for diastolic blood pressure from 0 to 0.108. CONCLUSION: The precise magnitude of between-cluster variation for a given measure can rarely be estimated in advance. Studies should be designed with reference to the overall distribution of ICCs and with attention to features that increase efficiency.

Cluster Analysis↗

Estimation of biogenic emissions with satellite-derived land use and land cover data for air quality modeling of Houston-Galveston ozone nonattainment area.

The Houston-Galveston Area (HGA) is one of the most severe ozone non-attainment regions in the US. To study the effectiveness of controlling anthropogenic emissions to mitigate regional ozone nonattainment problems, it is necessary to utilize adequate datasets describing the environmental conditions that influence the photochemical reactivity of the ambient atmosphere. Compared to the anthropogenic emissions from point and mobile sources, there are large uncertainties in the locations and amounts of biogenic emissions. For regional air quality modeling applications, biogenic emissions are not directly measured but are usually estimated with meteorological data such as photo-synthetically active solar radiation, surface temperature, land type, and vegetation database. In this paper, we characterize these meteorological input parameters and two different land use land cover datasets available for HGA: the conventional biogenic vegetation/land use data and satellite-derived high-resolution land cover data. We describe the procedures used for the estimation of biogenic emissions with the satellite derived land cover data and leaf mass density information. Air quality model simulations were performed using both the original and the new biogenic emissions estimates. The results showed that there were considerable uncertainties in biogenic emissions inputs. Subsequently, ozone predictions were affected up to 10 ppb, but the magnitudes and locations of peak ozone varied each day depending on the upwind or downwind positions of the biogenic emission sources relative to the anthropogenic NOx and VOC sources. Although the assessment had limitations such as heterogeneity in the spatial resolutions, the study highlighted the significance of biogenic emissions uncertainty on air quality predictions. However, the study did not allow extrapolation of the directional changes in air quality corresponding to the changes in LULC because the two datasets were based on vastly different LULC category definitions and uncertainties in the vegetation distributions.

Air Pollution↗

Partitioning protein structures into domains: why is it so difficult?

This analysis takes an in-depth look into the difficulties encountered by automatic methods for domain decomposition from three-dimensional structure. The analysis involves a multi-faceted set of criteria including the integrity of secondary structure elements, the tendency toward fragmentation of domains, domain boundary consistency and topology. The strength of the analysis comes from the use of a new comprehensive benchmark dataset, which is based on consensus among experts (CATH, SCOP and AUTHORS of the 3D structures) and covers 30 distinct architectures and 211 distinct topologies as defined by CATH. Furthermore, over 66% of the structures are multi-domain proteins; each domain combination occurring once per dataset. The performance of four automatic domain assignment methods, DomainParser, NCBI, PDP and PUU, is carefully analyzed using this broad spectrum of topology combinations and knowledge of rules and assumptions built into each algorithm. We conclude that it is practically impossible for an automatic method to achieve the level of performance of human experts. However, we propose specific improvements to automatic methods as well as broadening the concept of a structural domain. Such work is prerequisite for establishing improved approaches to domain recognition. (The benchmark dataset is available from http://pdomains.sdsc.edu).

Computational Biology↗

Topological models for the prediction of HIV-protease inhibitory activity of tetrahydropyrimidin-2-ones.

Relationship between the topological indices and HIV-protease inhibitory activity of tetrahydropyrimidine-2-ones has been investigated. Three topological indices, Wiener's index--a distance based topological descriptor, Zagreb group parameter--an adjacency based topological descriptor and eccentric connectivity index--an adjacency-cum-distance based topological descriptor were used for the present investigations. A dataset comprising of 80 substituted tetrahydropyrimidine-2-one analogues was selected for the present studies. The values of the Wiener's index, Zagreb group parameter and eccentric connectivity index for each of the 80 compounds comprising the dataset were computed using an in-house computer program. The dataset was divided randomly into training and test sets. Resultant data was analyzed and suitable models were developed after identifying the active ranges in the training set. Subsequently, a biological activity was assigned to each of the compound involved in the test set using these models, which was then compared with the reported HIV-protease inhibitory activity. Accuracy of prediction using these models was found to vary from a minimum of approximately 86% to a maximum of approximately 88%.

HIV Protease↗

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease↗

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans↗

Spatiotemporal patterns of Rift Valley fever virus in Africa: a retrospective genomic epidemiology and phylodynamic modelling study.

BACKGROUND: Rift Valley fever virus (RVFV) is a mosquito-borne zoonotic pathogen causing outbreaks in humans and ruminants across Africa and the Arabian Peninsula. Originally restricted to the Great Rift Valley, RVFV has expanded geographically, prompting its classification by WHO as a pathogen of pandemic potential. We investigated the evolutionary and spatial dynamics of RVFV across Africa. METHODS: We used genomic data generated at the International Livestock Research Institute Nairobi genomic laboratory (BioProject PRJNA1106221) and combined with publicly available datasets retrieved from the National Center for Biotechnology (NCBI) GenBank nucleotide database. In retrieving RVFV genome sequences from the NCBI GenBank, we applied the search terms "Rift Valley fever virus segment L AND 6404[SLEN]", "Rift Valley fever virus segment M AND 3885[SLEN]", and "Rift Valley fever virus segment S AND 1520:1690[SLEN]" for L (Large), M (Medium), and S (Small) segments, respectively. For sequences without additional spatiotemporal information, we searched PubMed to extract the associated sequence metadata. We performed molecular clock analysis, phylogenetic inference, phylodynamic modelling (continuous phylogeographic reconstruction), and landscape phylogeography on the three RVFV genome segments (L, M, and S). We aimed to assess evolutionary rates, dispersal patterns, and environmental drivers. Focus was placed on lineage C, the most widely distributed variant. FINDINGS: The global dataset used in this study consisted of large (n=236), medium (n=237), and small (n=247), which were further filtered to exclude potential reassortants and vaccine strains. Genome sequences retrieved from NCBI GenBank database comprised large (n=180), medium (n=184), and small (n=202). The genome sequences from retrospective human and livestock isolates comprised large (n=56), medium (n=53), and small (n=45) collected in Burundi (2018), Kenya (2007, 2018, 2019, 2021, and 2022), and Rwanda (2018 and 2022). Our dataset revealed that RVFV exhibited low overall genetic diversity. Lineage C, however, showed evidence of active evolution, with substitution rates ranging from 3&#xb7;58&#x2009;&#xd7;&#x2009;10-4 to 9&#xb7;76&#x2009;&#xd7;&#x2009;10-4 substitutions per site per year. This lineage probably originated in Zimbabwe in the mid-1970s and has since expanded across eastern and southern Africa. Phylogeographic reconstructions revealed rapid spread, with diffusion coefficients exceeding 50&#x2009;000 km2 per year. INTERPRETATION: Lineage C appears capable of establishing endemic transmission in new regions, with ongoing diversification observed during interepidemic periods. These observations reinforce the value of continuous genomic surveillance, particularly during cryptic transmission phases when adaptive mutations might emerge. Although further evidence is needed, observed trends in climate variability and land-use change point to the potential benefit of targeted surveillance in settings that could be at increased risk, including urban centres and wetlands. FUNDING: This work was supported by the German Federal Ministry for Economic Cooperation and Development, the Rockefeller Foundation, and the Africa Centres for Disease Control and Prevention.

Rift Valley fever virus↗

Systematic tuning of parameters in support vector clustering.

Clustering algorithms divide a set of observations into groups so that members of the same group share common features. In most of the algorithms, tunable parameters are set arbitrarily or by trial and error, resulting in less than optimal clustering. This paper presents a global optimization strategy for the systematic and optimal selection of parameter values associated with a clustering method. In the process, a performance criterion for the optimization model is proposed and benchmarked against popular performance criteria from the literature (namely, the Silhouette coefficient, Dunn's index, and Davies-Bouldin index). The tuning strategy is illustrated using the support vector clustering (SVC) algorithm and simulated annealing. In order to reduce the computational burden, the paper also proposes an alternative to the adjacency matrix method (used for the assignment of cluster labels), namely the contour plotting approach. Datasets tested include the iris and the thyroid datasets from the UCI repository, as well as lymphoma and breast cancer data. The optimal tuning parameters are determined efficiently, while the contour plotting approach leads to significant reductions in computational effort (CPU time) especially for large datasets. The performance criteria comparisons indicate mixed results. Specifically, the Silhouette coefficient and the Davies-Bouldin index perform better, while the Dunn's index is worse on average than the proposed performance index.

Algorithms↗

A rapid method for microarray cross platform comparisons using gene expression signatures.

Microarray technology has become highly valuable for identifying complex changes in global gene expression patterns. The inevitable use of a variety of different platforms has compounded the difficulty of effectively comparing data between projects, laboratories, and public access databases. The need for consistent, believable results across platforms is fundamental and methods for comparing results across platforms should be as straightforward as possible. We present the results of a study comparing three major, commercially available, microarray platforms (Affymetrix, Agilent, and Illumina). Concordance estimates between platforms was based on mapping of probes to Human Gene Organization (HUGO) gene names. Appropriate data normalization procedures were applied to each dataset followed by the generation of lists of regulated genes using a common significance threshold for all three platforms. As expected, concordance measured by directly comparing gene lists was relatively low (an average 22.8% for all platforms across all possible comparisons). However, when statistical tests (gene set enrichment analysis--GSEA, parametric analysis of gene enrichment--PAGE) which align gene lists with continuous measures of differential gene expression were applied to the cross platform datasets using significant gene lists to poll entire datasets, the relatedness of the results from all three platforms was specific, obvious, and profound.

Animals↗

Data-partitioning using the Hilbert space filling curves: effect on the speed of convergence of Fuzzy ARTMAP for large database problems.

The Fuzzy ARTMAP algorithm has been proven to be one of the premier neural network architectures for classification problems. One of the properties of Fuzzy ARTMAP, which can be both an asset and a liability, is its capacity to produce new nodes (templates) on demand to represent classification categories. This property allows Fuzzy ARTMAP to automatically adapt to the database without having to a priori specify its network size. On the other hand, it has the undesirable side effect that large databases might produce a large network size (node proliferation) that can dramatically slow down the training speed of the algorithm. To address the slow convergence speed of Fuzzy ARTMAP for large database problems, we propose the use of space-filling curves, specifically the Hilbert space-filling curves (HSFC). Hilbert space-filling curves allow us to divide the problem into smaller sub-problems, each focusing on a smaller than the original dataset. For learning each partition of data, a different Fuzzy ARTMAP network is used. Through this divide-and-conquer approach we are avoiding the node proliferation problem, and consequently we speedup Fuzzy ARTMAP's training. Results have been produced for a two-class, 16-dimensional Gaussian data, and on the Forest database, available at the UCI repository. Our results indicate that the Hilbert space-filling curve approach reduces the time that it takes to train Fuzzy ARTMAP without affecting the generalization performance attained by Fuzzy ARTMAP trained on the original large dataset. Given that the resulting smaller datasets that the HSFC approach produces can independently be learned by different Fuzzy ARTMAP networks, we have also implemented and tested a parallel implementation of this approach on a Beowulf cluster of workstations that further speeds up Fuzzy ARTMAP's convergence to a solution for large database problems.

Algorithms↗

Graph kernels for chemical informatics.

Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MinMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed.

Anticarcinogenic Agents↗

Comparative analysis of convolutional neural network models for the histopathological differentiation of acinic cell carcinoma and secretory carcinoma.

OBJECTIVE: Although artificial intelligence tools show promise for enhancing the diagnosis of head and neck lesions, few studies have tested these resources for the microscopic diagnosis of salivary gland tumors. Specifically, the microscopic differentiation between acinic cell carcinoma and secretory carcinoma has never been addressed in this context. Therefore, this exploratory study aimed to comparatively evaluate the feasibility of applying convolutional neural networks for the microscopic differentiation between acinic cell carcinomas and secretory carcinomas. METHODS: A cross-sectional study using whole-slide images from 46 patients with acinic cell carcinoma (n = 26) or secretory carcinoma (n = 20) was conducted. Eight CNNs (ResNet-50, InceptionV3, VGG16, Xception, MobileNet, DenseNet121, EfficientNetB0, and EfficientNetV2B0) were trained and evaluated for accuracy, sensitivity, specificity, F1-score, and AUC. Performance was measured in training, validation, and test subsets. Accuracy and loss curves were also presented. RESULTS: InceptionV3 demonstrated the best overall performance, with the lowest loss (1.39), highest accuracy (0.81), sensitivity (0.90), and F1-score (0.81). VGG16 achieved the highest AUC (0.86) and precision (0.77). DenseNet121 showed the lowest performance in terms of accuracy (0.65) and F1-Score (0.52), but the highest specificity (0.85). CONCLUSION: This proof-of-concept study suggests that convolutional neural networks may be feasible tools to support the microscopic differentiation between acinic cell carcinoma and secretory carcinoma. The performance of these models critically depends on the size of the dataset and the quality of annotations. The findings should be interpreted cautiously given the limited dataset and potential sources of bias. Further validation with larger, multicenter datasets is needed before any clinical application can be considered.

Humans↗

Gamma histograms for radiotherapy plan evaluation.

BACKGROUND AND PURPOSE: The technique known as the 'gamma evaluation method' incorporates pass-fail criteria for both distance-to-agreement and dose difference analysis of 3D dose distributions and provides a numerical index (gamma) as a measure of the agreement between two datasets. As the gamma evaluation index is being adopted in more centres as part of treatment plan verification procedures for 2D and 3D dose maps, the development of methods capable of encapsulating the information provided by this technique is recommended. PATIENTS AND METHODS: In this work the concept of gamma index was extended to create gamma histograms (GH) in order to provide a measure of the agreement between two datasets in two or three dimensions. Gamma area histogram (GAH) and gamma volume histogram (GVH) graphs were produced using one or more 2D gamma maps generated for each slice of the irradiated volume. GHs were calculated for IMRT plans, evaluating the 3D dose distribution from a commercial treatment planning system (TPS) compared to a Monte Carlo (MC) calculation used as reference dataset. RESULTS: The extent of local anatomical inhomogenities in the plans under consideration was strongly correlated with the level of difference between reference and evaluated calculations. GHs provided an immediate visual representation of the proportion of the treated volume that fulfilled the gamma criterion and offered a concise method for comparative numerical evaluation of dose distributions. CONCLUSIONS: We have introduced the concept of GHs and investigated its applications to the evaluation and verification of IMRT plans. The gamma histogram concept set out in this paper can provide a valuable technique for quantitative comparison of dose distributions and could be applied as a tool for the quality assurance of treatment planning systems.

Head and Neck Neoplasms↗

Comparison of helical, maximum intensity projection (MIP), and averaged intensity (AI) 4D CT imaging for stereotactic body radiation therapy (SBRT) planning in lung cancer.

BACKGROUND AND PURPOSE: To compare helical, MIP and AI 4D CT imaging, for the purpose of determining the best CT-based volume definition method for encompassing the mobile gross tumor volume (mGTV) within the planning target volume (PTV) for stereotactic body radiation therapy (SBRT) in stage I lung cancer. MATERIALS AND METHODS: Twenty patients with medically inoperable peripheral stage I lung cancer were planned for SBRT. Free-breathing helical and 4D image datasets were obtained for each patient. Two composite images, the MIP and AI, were automatically generated from the 4D image datasets. The mGTV contours were delineated for the MIP, AI and helical image datasets for each patient. The volume for each was calculated and compared using analysis of variance and the Wilcoxon rank test. A spatial analysis for comparing center of mass (COM) (i.e. isocenter) coordinates for each imaging method was also performed using multivariate analysis of variance. RESULTS: The MIP-defined mGTVs were significantly larger than both the helical- (p=0.001) and AI-defined mGTVs (p=0.012). A comparison of COM coordinates demonstrated no significant spatial difference in the x-, y-, and z-coordinates for each tumor as determined by helical, MIP, or AI imaging methods. CONCLUSIONS: In order to incorporate the extent of tumor motion from breathing during SBRT, MIP is superior to either helical or AI images for defining the mGTV. The spatial isocenter coordinates for each tumor were not altered significantly by the imaging methods.

Humans↗

Creating and using the urgent metadata catalogue and thesaurus.

The Urban Regeneration and the Environment Research Programme (URGENT) required a system for cataloguing its datasets and enabling its scientific community to discover what data were available to it. This community was multidisciplinary in nature and therefore needed a range of facilities for searching. Of particular importance were facilities to help those unfamiliar with specialist terminology. To meet these needs, four applications were designed and developed: a Metadata Capture Tool for describing datasets in compliance with the National Geospatial Data Framework (NGDF) standard, a Term Entry Tool for creating an ISO compliant thesaurus, a Thesaurus Builder for merging thesauri and a Search Tool. To encourage users to help in cataloguing data, the capture tools were written as stand alone applications, which users could keep and use to build their own metadatabases. The tools contained export and import facilities that allowed the URGENT Data Centre to build a central database and publish it upon the web. During the development work, it was found necessary to extend the NGDF standard as it could not adequately describe time variant or 3-D atmospheric datasets. The four applications met their design objectives. However, a number of ergonomic issues will need to be addressed if the system is to meet the needs of the much larger up coming programmes. The main challenges will be moving from the NGDF standard to the ISO standard, hence bringing the work into line with the recommendations of the INSPIRE Project, and merging the metadatabase with the scientific database, which enable metadata maintenance to be semi-automated.

Cataloging↗

How to analyze and understand the human immune system.

To enhance our understanding of the pathogenesis of diseases, including rheumatic diseases, and to improve disease control, it is essential to attain a thorough understanding of the human immune system, alongside mouse immunology. Historically, the investigation of the human immune system has posed significant challenges due to methodological limitations. Nonetheless, recent advancements in genomic studies of multifactorial diseases have elucidated that numerous risk-associated genetic variants affecting quantitative differences in cell-specific gene expression. In light of these findings, we are currently examining individual genetic variations in both healthy individuals and patients, as well as categorizing cells into distinct subsets in order to construct a comprehensive dataset concerning the human immune system. This is accomplished by combining data on gene expression, factors influencing the expression mechanisms, protein expression, metabolomics, and environmental variables pertinent to immune functionality-such as gut microbiota. These datasets will facilitate the comprehensive characterization of the human immune system. Using these datasets and through the integrative analyses of data related to risk genetic variations and gene expression profiles of each disease and individual, we anticipate uncovering novel insights into the human immune system, the heterogeneity of diseases, immune function mechanisms, and their regulatory strategies that may not be achievable through murine models.

Humans↗