Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Psychometric properties of an index of emotional distress in the U.S. National Comorbidity Survey.

BACKGROUND: The National Comorbidity Survey (NCS; Kessler et al. 1994) was a nationwide household survey of the U.S. population designed to produce data on the prevalence and correlates of psychiatric disorders. The NCS dataset is now in public-use format and continues to be widely used for ongoing research efforts. The NCS dataset included a set of 14 items that have face validity as a measure of current emotional distress (depression and anxiety) and could serve as a potentially useful continuous measure of psychological distress. However, there have been no published studies on its psychometric properties and this measure has not yet been utilized by researchers using the NCS dataset. This paper provides an evaluation of the psychometric properties of the NCS Distress Index. METHOD: The NCS Part II public-use dataset (N = 5877) was used. Detailed diagnostic information was collected along with 14 items assessing current psychological distress and measures of Neuroticism and Openness to Experience. RESULTS: The NCS Distress Index was found to be internally consistent (Alpha = 0.92) and a series of principal-components analyses demonstrated that the measure is most accurately conceptualized as a single-factor measure of general distress. The construct validity of the Distress Index was supported by its associations with the measures of Neuroticism and Openness to Experience. A series of comparisons between diagnostic groups also supported the construct validity of the measure. For example, those with disorders characterized by depressed mood and worry scored higher on the Distress Index than those with disorders characterized by fear and hyperarousal. CONCLUSIONS: The NCS Distress Index is a psychometrically sound measure of current emotional distress. Future studies utilizing the NCS public-use dataset could potentially benefit from the inclusion of this measure in addition to more commonly investigated categorical variables such as diagnosable disorders.

Adolescent↗

Detection of osteoporotic vertebral fractures using multidetector CT.

INTRODUCTION: Goals were to compare the performance of lateral radiographs and sagittal reformations (SR) of axial computed tomography (CT) datasets in identification of osteoporotic vertebral fractures and to assess for optimal slice thickness in axial CT datasets needed for reliable classification of these fractures. METHODS: Sixty-five vertebrae were harvested from 21 human cadaver spines and examined with a 64-row multidetector CT scanner. Axial images were acquired with a slice thickness of 0.6, 1, 2, 3 and 5 mm and SR were obtained using these datasets. In addition, specimens were radiographed in antero-posterior and lateral orientation. Vertebrae visualized in the different image datasets were separately graded by four radiologists according to the spinal fracture index (SFI) classification. Fracture status determined in a consensus reading of interactive reformations of the 0.6-mm CT dataset in all three dimensions served as a standard of reference in combination with pathological examinations. RESULTS: The average agreement for the 0.6-mm SR obtained between each radiologist and standard of reference for the grading of the fractures was very good (kappa=0.81). It was good for the 1-, 2- and 3-mm SR (kappa=0.70, 0.69 and 0.64), but only moderate for the radiographs (kappa=0.52), and fair for the 5-mm SR (kappa=0.33). When focusing only on detection of fractures, independent of the grading, all kappa values improved by about 0.15, resulting in excellent values for the 0.6-mm through 3-mm SR (0.95<kappa<0.79) and good values for the radiographs (kappa=0.72). Ninety-five percent of the fractures could be identified using the 1-mm SR, but 18% of the fractures were missed on the radiographs. CONCLUSIONS: Sagittal CT reformations could more accurately assess vertebral fractures than standard radiographs. But for reliable detection of these fractures, SR derived from axial images with a slice thickness of 3 mm or less are required. The thinnest available axial slice thickness performed best in fracture grading.

Aged, 80 and over↗

Application of the character compatibility approach to generalized molecular sequence data: branching order of the proteobacterial subdivisions.

The character compatibility approach, which removes all homoplasic characters and involves finding the largest clique of compatible characters in a dataset, in principle, provides a powerful means for obtaining correct topology in difficult to resolve cases. However, the usefulness of this approach to generalized molecular sequence data for phylogeny determination has not been studied in the past. We have used this approach to determine the topology of 23 proteobacterial species (6 each of alpha-, beta- and gamma-, 3 delta-, and 2 epsilon-proteobacteria) using sequence data for 10 conserved proteins (Hsp60, Hsp70, EF-Tu, EF-G, alanyl-tRNA synthetase, RecA, GyrA, GyrB, RpoB and RpoC). All sites in the sequence alignments of these proteins where only two amino acids were found, with each amino acid present in at least two species, were selected. Mutual compatibility determination on these binary state sites was carried out by two means. In one case, all of these sites were combined into a large dataset (Set A; 957 characters) prior to compatibility analysis. In the second case, compatibility analysis was carried out on characters from individual proteins and all compatible sites were combined into a large dataset (Set B; 398 characters) for further studies. Upon compatibility analyses, the largest cliques that were obtained from Sets A and B consisted of 337 and 323 compatible characters, respectively. In these cliques, all proteobacterial subgroups were clearly distinguished and branching orders of most of the species were also resolved. The epsilon-proteobacteria exhibited the earliest branching, whereas the beta- and gamma-subgroups were found to have emerged last. The relative placement of the alpha- and delta-subgroups, however, was not resolved. The topology of these species was also determined based on 16S rRNA sequences and a concatenated dataset of sequences for all 10 proteins by means of neighbor-joining, maximum likelihood, and maximum parsimony methods. In the protein trees, all proteobacterial groups were reliably resolved and they branched in the following order: (epsilon(delta(alpha(beta,gamma)))). However, in the rRNA trees, the gamma- and beta-subgroups exhibited polyphyletic branching and many internal nodes were not resolved. These results indicate that the character compatibility analysis using generalized molecular sequence data provides a powerful means for evolutionary studies. Based on molecular sequences, it should be possible to obtain very large datasets of compatible characters that should prove very helpful in clarifying difficult to resolve phylogenetic relationships.

Amino Acid Sequence↗

Characterization of clustered microcalcifications in digitized mammograms using neural networks and support vector machines.

OBJECTIVE: Detection and characterization of microcalcification clusters in mammograms is vital in daily clinical practice. The scope of this work is to present a novel computer-based automated method for the characterization of microcalcification clusters in digitized mammograms. METHODS AND MATERIAL: The proposed method has been implemented in three stages: (a) the cluster detection stage to identify clusters of microcalcifications, (b) the feature extraction stage to compute the important features of each cluster and (c) the classification stage, which provides with the final characterization. In the classification stage, a rule-based system, an artificial neural network (ANN) and a support vector machine (SVM) have been implemented and evaluated using receiver operating characteristic (ROC) analysis. The proposed method was evaluated using the Nijmegen and Mammographic Image Analysis Society (MIAS) mammographic databases. The original feature set was enhanced by the addition of four rule-based features. RESULTS AND CONCLUSIONS: In the case of Nijmegen dataset, the performance of the SVM was Az=0.79 and 0.77 for the original and enhanced feature set, respectively, while for the MIAS dataset the corresponding characterization scores were Az=0.81 and 0.80. Utilizing neural network classification methodology, the corresponding performance for the Nijmegen dataset was Az=0.70 and 0.76 while for the MIAS dataset it was Az=0.73 and 0.78. Although the obtained high classification performance can be successfully applied to microcalcification clusters characterization, further studies must be carried out for the clinical evaluation of the system using larger datasets. The use of additional features originating either from the image itself (such as cluster location and orientation) or from the patient data may further improve the diagnostic value of the system.

Breast Diseases↗

The classification of cancer based on DNA microarray data that uses diverse ensemble genetic programming.

OBJECT: The classification of cancer based on gene expression data is one of the most important procedures in bioinformatics. In order to obtain highly accurate results, ensemble approaches have been applied when classifying DNA microarray data. Diversity is very important in these ensemble approaches, but it is difficult to apply conventional diversity measures when there are only a few training samples available. Key issues that need to be addressed under such circumstances are the development of a new ensemble approach that can enhance the successful classification of these datasets. MATERIALS AND METHODS: An effective ensemble approach that does use diversity in genetic programming is proposed. This diversity is measured by comparing the structure of the classification rules instead of output-based diversity estimating. RESULTS: Experiments performed on common gene expression datasets (such as lymphoma cancer dataset, lung cancer dataset and ovarian cancer dataset) demonstrate the performance of the proposed method in relation to the conventional approaches. CONCLUSION: Diversity measured by comparing the structure of the classification rules obtained by genetic programming is useful to improve the performance of the ensemble classifier.

Artificial Intelligence↗

Amino acid propensities for secondary structures are influenced by the protein structural class.

Amino acid propensities for secondary structures were used since the 1970s, when Chou and Fasman evaluated them within datasets of few tens of proteins and developed a method to predict secondary structure of proteins, still in use despite prediction methods having evolved to very different approaches and higher reliability. Propensity for secondary structures represents an intrinsic property of amino acid, and it is used for generating new algorithms and prediction methods, therefore our work has been aimed to investigate what is the best protein dataset to evaluate the amino acid propensities, either larger but not homogeneous or smaller but homogeneous sets, i.e., all-alpha, all-beta, alpha-beta proteins. As a first analysis, we evaluated amino acid propensities for helix, beta-strand, and coil in more than 2000 proteins from the PDBselect dataset. With these propensities, secondary structure predictions performed with a method very similar to that of Chou and Fasman gave us results better than the original one, based on propensities derived from the few tens of X-ray protein structures available in the 1970s. In a refined analysis, we subdivided the PDBselect dataset of proteins in three secondary structural classes, i.e., all-alpha, all-beta, and alpha-beta proteins. For each class, the amino acid propensities for helix, beta-strand, and coil have been calculated and used to predict secondary structure elements for proteins belonging to the same class by using resubstitution and jackknife tests. This second round of predictions further improved the results of the first round. Therefore, amino acid propensities for secondary structures became more reliable depending on the degree of homogeneity of the protein dataset used to evaluate them. Indeed, our results indicate also that all algorithms using propensities for secondary structure can be still improved to obtain better predictive results.

Amino Acid Motifs↗

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine↗

Further support for the clades obtained by multiple molecular phylogenies in the acanthomorph bush.

Several recent molecular studies have begun to clarify the phylogeny of Acanthomorpha (Teleostei), a wide clade of teleost fishes. However, different molecular datasets do not agree on a single history of the taxa, probably because of marker-specific biases. The 'total-evidence' approach maximizes character congruence, but may be biased by a single robust, but non-phylogenetic constraint from one dataset. We have therefore taken the approach to analyse also each dataset separately prior to their combination, and detect repeated groups: signal common to markers is more probably a reflection of shared ancestry than marker-specific signal. Partial sequences (678+527 base pairs) of exons of the MLL gene (Mixed Lineage Leukaemia-like) gene were used, as well as the datasets of Chen et al. (ribosomal 28S, rhodopsin gene, mitochondrial 12S and 16S). Most of the repeated clades of Chen et al. are supported by the new dataset. Some new groups were repeatedly found: a Scarus-Labrus group (clade M), the presence of Gasterosteidae as a sister taxon or within the clade Zoarcoidei-Cottoidei (clade Is), Polymixia as a sister-group to the clade Zeoidei-Gadiformes (clade O), the clade Q grouping Mugiloidei, Cichlidae, Atherinomorpha, Blennioidei and Gobiesocoidei; and the interesting clade N, reducing potential sister-groups to Tetraodontiformes to either Caproidei, Lophiiformes, Acanthuroidei, Drepanidae, Chaetodontidae, and Pomacanthidae.

Animals↗

GEMS: a system for automated cancer diagnosis and biomarker discovery from microarray gene expression data.

The success of treatment of patients with cancer depends on establishing an accurate diagnosis. To this end, we have built a system called GEMS (gene expression model selector) for the automated development and evaluation of high-quality cancer diagnostic models and biomarker discovery from microarray gene expression data. In order to determine and equip the system with the best performing diagnostic methodologies in this domain, we first conducted a comprehensive evaluation of classification algorithms using 11 cancer microarray datasets. In this paper we present a preliminary evaluation of the system with five new datasets. The performance of the models produced automatically by GEMS is comparable or better than the results obtained by human analysts. Additionally, we performed a cross-dataset evaluation of the system. This involved using a dataset to build a diagnostic model and to estimate its future performance, then applying this model and evaluating its performance on a different dataset. We found that models produced by GEMS indeed perform well in independent samples and, furthermore, the cross-validation performance estimates output by the system approximate well the error obtained by the independent validation. GEMS is freely available for download for non-commercial use from http://www.gems-system.org.

Algorithms↗

SAR modeling of genotoxic phenomena: the consequence on predictive performance of deviation from a unity ratio of genotoxicants/non-genotoxicants.

SAR approaches to the study of genotoxic phenomena are finding increased applications. However, the data being modeled are frequently not considered optimal due to the small size of the dataset and an uneven distribution of genotoxicants and non-genotoxicants in the dataset. The effects of such imbalances on the performance of one SAR approach were investigated with respect to the modeling of the induction of unscheduled DNA synthesis in rat hepotocytes and of sister chromatid exchanges and chromosomal aberrations in cultural CHO cells. The analyses indicate that if genotoxicants exceed non-genotoxicants, the performance of the SAR model can be improved if the dataset is supplemented with physiological chemicals which are assumed to be non-genotoxicants. On the other hand, if non-genotoxicants exceed genotoxicants, it was found that the predictive performance of the resulting SAR model is not improved by removal of genotoxicants from the dataset to achieve a ratio of genotoxicants/genotoxicants of unity. Overall, the present analyses did not result in the development of SAR models of greatly increased predictivity. Conceivably, for the particular datasets and SAR paradigm the limit of predictivity has been reached. The possibility of investigating the use of a "battery" of SAR paradigms should be considered.

Models, Biological↗

Relation between the generation time and the lag time of bacterial growth kinetics.

In predictive microbiology, the relation between the lag time (Lag) and the generation time (Tg) is commonly assumed to be proportional, as long as the pre-incubation environmental conditions remain constant. This relation was statistically examined in nine published datasets. For every dataset, it was roughly proportional. However, a more advanced study showed that the ratio Lag/Tg was not totally independent of the environmental conditions. In particular, a significant negative effect of the pH on this ratio was observed in five of the nine datasets. For modeling the environmental dependence of microbial growth parameters, some authors independently deal with Lag and Tg. Other authors only model the environmental dependence of Tg, assuming Lag/Tg to be constant. These two modeling methods were statistically compared for the nine datasets under study. Results differed from one dataset to another. For some, the model developed with a constant ratio Lag/Tg sufficed to describe the data, whereas for the others, an independent modeling of Lag and Tg was more satisfactory.

Bacteria↗

Prediction of beta-turns with learning machines.

The support vector machine approach was introduced to predict the beta-turns in proteins. The overall self-consistency rate by the re-substitution test for the training or learning dataset reached 100%. Both the training dataset and independent testing dataset were taken from Chou [J. Pept. Res. 49 (1997) 120]. The success prediction rates by the jackknife test for the beta-turn subset of 455 tetrapeptides and non-beta-turn subset of 3807 tetrapeptides in the training dataset were 58.1 and 98.4%, respectively. The success rates with the independent dataset test for the beta-turn subset of 110 tetrapeptides and non-beta-turn subset of 30,231 tetrapeptides were 69.1 and 97.3%, respectively. The results obtained from this study support the conclusion that the residue-coupled effect along a tetrapeptide is important for the formation of a beta-turn.

Artificial Intelligence↗

Comparing performance of multinomial logistic regression and discriminant analysis for monitoring access to care for acute myocardial infarction.

One way to monitor patient access to emergent health care services is to use patient characteristics to predict arrival time at the hospital after onset of symptoms. This predicted arrival time can then be compared with actual arrival time to allow monitoring of access to services. Predicted arrival time could also be used to estimate potential effects of changes in health care service availability, such as closure of an emergency department or an acute care hospital. Our goal was to determine the best statistical method for prediction of arrival intervals for patients with acute myocardial infarction (AMI) symptoms. We compared the performance of multinomial logistic regression (MLR) and discriminant analysis (DA) models. Models for MLR and DA were developed using a dataset of 3,566 male veterans hospitalized with AMI in 81 VA Medical Centers in 1994-1995 throughout the United States. The dataset was randomly divided into a training set (n = 1,846) and a test set (n = 1,720). Arrival times were grouped into three intervals on the basis of treatment considerations: <6 hours, 6-12 hours, and >12 hours. One model for MLR and two models for DA were developed using the training dataset. One DA model had equal prior probabilities, and one DA model had proportional prior probabilities. Predictive performance of the models was compared using the test (n = 1,720) dataset. Using the test dataset, the proportions of patients in the three arrival time groups were 60.9% for <6 hours, 10.3% for 6-12 hours, and 28.8% for >12 hours after symptom onset. Whereas the overall predictive performance by MLR and DA with proportional priors was higher, the DA models with equal priors performed much better in the smaller groups. Correct classifications were 62.6% by MLR, 62.4% by DA using proportional prior probabilities, and 48.1% using equal prior probabilities of the groups. The misclassifications by MLR for the three groups were 9.5%, 100.0%, 74.2% for each time interval, respectively. Misclassifications by DA models were 9.8%, 100.0%, and 74.4% for the model with proportional priors and 47.6%, 79.5%, and 51.0% for the model with equal priors. The choice of MLR or DA with proportional priors, or DA with equal priors for monitoring time intervals of predicted hospital arrival time for a population should depend on the consequences of misclassification errors.

Aged↗

A comparison between two neural network rule extraction techniques for the diagnosis of hepatobiliary disorders.

Neural networks have been widely used as tools for prediction in medicine. We expect to see even more applications of neural networks for medical diagnosis as recently developed neural network rule extraction algorithms make it possible for the decision process of a trained network to be expressed as classification rules. These rules are more comprehensible to a human user than the classification process of the networks which involves complex nonlinear mapping of the input data. This paper reports the results from two neural network rule extraction techniques, NeuroLinear and NeuroRule applied to the diagnosis of hepatobiliary disorders. The dataset consists of nine measurements collected from patients in a Japanese hospital and these measurements have continuous values. NeuroLinear generates piece-wise linear discriminant functions for this dataset. The continuous measurements have previously been discretized by domain experts. NeuroRule is applied to the discretized dataset to generate symbolic classification rules. We compare the rules generated by the two techniques and find that the rules generated by NeuroLinear from the original continuously valued dataset to be slightly more accurate and more concise than the rules generated by NeuroRule from the discretized dataset.

Algorithms↗

Molecular phylogenetics of myliobatiform fishes (Chondrichthyes: Myliobatiformes), with comments on the effects of missing data on parsimony and likelihood.

Mitochondrial DNA sequences from the 12S rRNA gene, four tRNA genes, and a portion of two protein coding genes were used to investigate the relationship of myliobatoid genera. In addition, we conducted an investigation of the sister group to the freshwater stingrays by sampling additional DNA sequences from GenBank. Consequently, two datasets were used to examine myliobatoid relationships. The first consisted of the genes sequenced in this study. The second dataset was compiled by combining the first dataset with cytochrome b sequences from GenBank. The second dataset, however, included a number of missing characters due to differences in sampling. The effect of the missing characters on both maximum parsimony and maximum likelihood analysis was investigated by conducting a simulation study. Results of the simulation study indicated that maximum likelihood was not sensitive to the missing data, whereas the accuracy of maximum parsimony analysis was expected to decrease. Phylogenetic analysis of this group had several areas concordant with morphological studies, however, the analysis also revealed two novel relationships. In addition, placement of two taxa (Gymnura and Himantura) were dependent both on the dataset and analytical method used.

Animals↗

Application of three-way principal component analysis to the evaluation of two-dimensional maps in proteomics.

Three-way PCA has been applied to proteomic pattern images to identify the classes of samples present in the dataset. The developed method has been applied to two different datasets: a rat sera dataset, constituted by five samples of healthy Wistar rat sera and five samples of nicotine-treated Wistar rat sera; a human lymph-node dataset constituted by four healthy lymph-nodes and four lymph-nodes affected by a non-Hodgkin's lymphoma. The method proved to be successful in the identification of the classes of samples present in both of the groups of 2D-PAGE images, and it allowed us to identify the regions of the two-dimensional maps responsible for the differences occurring between the classes for both rat sera and human lymph-nodes datasets.

Algorithms↗

Analytical reproducibility in (1)H NMR-based metabonomic urinalysis.

Metabonomic analysis of biofluids and tissues utilizing high-resolution NMR spectroscopy and chemometric techniques has proven valuable in characterizing the biochemical response to toxicity for many xenobiotics. To assess the analytical reproducibility of metabonomic protocols, sample preparation and NMR data acquisition were performed at two sites (one using a 500 MHz and the other using a 600 MHz system) using two identical (split) sets of urine samples from an 8-day acute study of hydrazine toxicity in the rat. Despite the difference in spectrometer operating frequency, both datasets were extremely similar when analyzed using principal components analysis (PCA) and gave near-identical descriptions of the metabolic responses to hydrazine treatment. The main consistent difference between the datasets was related to the efficiency of water resonance suppression in the spectra. In a 4-PC model of both datasets combined, describing all systematic dose- and time-related variation (88% of the total variation), differences between the two datasets accounted for only 3% of the total modeled variance compared to ca. 15% for normal physiological (pre-dose) variation. Furthermore, <3% of spectra displayed distinct inter-site differences, and these were clearly identified as outliers in their respective dose-group PCA models. No samples produced clear outliers in both datasets, suggesting that the outliers observed did not reflect an unusual sample composition, but rather sporadic differences in sample preparation leading to, for example, very dilute samples. Estimations of the relative concentrations of citrate, hippurate, and taurine were in >95% correlation (r(2)) between sites, with an analytical error comparable to normal physiological variation in concentration (4-8%). The excellent analytical reproducibility and robustness of metabonomic techniques demonstrated here are highly competitive compared to the best proteomic analyses and are in significant contrast to genomic microarray platforms, both of which are complementary techniques for predictive and mechanistic toxicology. These results have implications for the quantitative interpretation of metabonomic data, and the establishment of quality control criteria for both regulatory agencies and for integrating data obtained at different sites.

Animals↗

Population of the HLA ligand database.

We have established an HLA ligand database to provide scientists and clinicians with access to Major Histocompatibility Complex (MHC) class I and II motif and ligand data. The HLA Ligand Database is available on the world wide web at http://hlaligand.ouhsc.edu and contains ligands that have been published in peer-reviewed journals. HLA peptide datasets prove useful in several areas: ligands are important as targets for various immune responses while algorithms built upon ligand datasets allow identification of new peptides without time-consuming experimental procedures. A review of the HLA class I ligands in the database identifies strengths and deficiencies in the database and, therefore, the utility of the dataset for identifying new peptides. For instance, 212 HLA-A phenotypes exist of which 23 have a motif determined and 43 have peptides characterized. In terms of number of ligands, HLA-A*0201 has 258 characterized ligands, A*1101 has 25 peptides, while the remaining two-thirds of the HLA-A phenotypes have less than 10 associated peptide sequences. Characterization of ligands and motifs remains roughly the same at the HLA-B locus while the peptides of the HLA-C locus tend to be less characterized. These data show that 74% of HLA class I molecules do not have ligands represented in the database and thus algorithms based on the dataset could not predict ligands for a majority of the US population. Building upon this dataset and knowledge of HLA allelic frequencies, it is possible to plan a systematic expansion of the HLA class I ligand database to better identify ligands useful throughout the population.

Databases, Protein↗