Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

The SAGES Bariatric Surgery Outcome Initiative.

BACKGROUND: The recent initiative for identifying centers of excellence in bariatric surgery calls for documentation of surgical outcomes. The SAGES Outcomes Initiative is a national database introduced in 1999 as a method for surgeons to accumulate and compare their data with summary national data. A bariatric-specific dataset was established later in 2001. The aim of this study was to compare the outcomes of bariatric surgery from the Society of American Gastrointestinal Endoscopic Surgeons' (SAGES) bariatric database with data derived from a national administrative database of academic centers. METHODS: Between 2001 and 2004, 24 surgeons with 1,954 patients participated in the SAGES Bariatric Outcome Initiative, and 97 institutions with 42,847 patients participated in the University HealthSystem Consortium (UHC) database. Only 7 of the 24 surgeons participating in the SAGES Bariatric Outcome Initiative submitted more than 50 cases. The main outcome measures included demographics, comorbidities, type of bariatric procedure, operative time, length of hospital stay, short- and long-term complications, mortality, and weight loss. RESULTS: Both datasets were comparable for gender. Roux-en-Y gastric bypass had been performed for 88% of the patients in the SAGES database and 96% of the patients in the UHC database. Associated comorbidities were similar between the two groups except for a higher rate of hyperlipidemia for the patients in the SAGES database. The SAGES database contains more bariatric-specific information such as body mass index, operative time, blood loss, bariatric-specific complications, long-term complications, and weight loss data than the UHC database. According to the available data, no statistically significant differences exist between the two datasets in terms of perioperative complications and mortality. CONCLUSIONS: The SAGES Bariatric Outcome Initiative provides valuable bariatric-specific data not currently available in an administrative database that may be useful for benchmarking purposes. However, this database is currently underutilized.

Adult↗

Investigation of seven proposed regions of linkage in multiple sclerosis: an American and French collaborative study.

Multiple sclerosis (MS) is a demyelinating autoimmune disease with a strong yet complex genetic component. To date only the HLA-DR locus, and specifically the HLA-DR15 allele, has been identified and confirmed as influencing the risk of developing MS. Genomic screens on several datasets have been performed and have identified several chromosomal regions with interesting results, but none have yet been confirmed. We tested seven of the most-promising regions (on chromosomes 1p, 2p, 3p, 3q, 5q, 19q, and Xp) identified from several genomic screens in a dataset of 98 multiplex MS families from the United States and 90 multiplex MS families from France. The results did not confirm linkage to 2p, 3q, 5q, or Xp in the overall dataset, or in subsets defined by geographic origin or HLA-DR15 status. Regions on 1p34, 3p14, and 19q13 produced lod scores >0.90 in at least one subset of the data, suggesting that these regions should be examined in more detail.

Chromosomes, Human, Pair 1↗

Notch pathway defines an aggressive and immune-suppressive phenotype associated with checkpoint inhibitor resistance in pan-gastrointestinal adenocarcinomas.

The Notch pathway regulates the homeostasis and tumorigenesis of gastrointestinal epithelium. Given its roles in cancer stem cell capacity and cancer immunity, we hypothesized that Notch activation can predict poor prognosis and resistance to immune checkpoint inhibitors (ICIs) in gastrointestinal adenocarcinoma (GIAC). The mRNA expression and genomic alterations of Notch pathway were characterized in esophagus (ESAD), stomach (STAD), colon (COAD), or rectum (READ) adenocarcinomas from The Cancer Genome Atlas (TCGA) dataset. The prognostic model (mRNA-score) was constructed using the TCGA dataset (the training set) and was validated in 3 independent sets (GSE19417 [ESAD], GSE84437 [STAD], and GSE40967 [COAD]). The associations of the mRNA-score with drug sensitivity, immune cell infiltration, and immunotherapy efficacy were, respectively, analyzed using the Genomics of Drug Sensitivity in Cancer (GDSC) database, the TCGA dataset, and multiple clinical cohorts including GSE165252, PRJEB25780, IMvigor210, and CheckMate-009/010/025. Notch pathway genes exhibited conserved genomic/transcriptomic features across four GIAC subtypes. Three pan-GIAC clusters were determined by unsupervised clustering, and the cluster with higher expression of the Notch pathway genes had shorter overall survival (OS), immunosuppressive microenvironment, and higher scores of the signatures concerning angiogenesis, cell cycle, PI3K-AKT-mTOR, TGF-&#x3b2;, glycolysis, etc. A prognostic algorithm (mRNA-score) was constructed, which was correlated with poor OS in the training set (TCGA, P&#x2009;<&#x2009;0.001) and three validation sets (GSE19417, P&#x2009;=&#x2009;0.025; GSE84437, P&#x2009;=&#x2009;0.001, GSE40967, P&#x2009;=&#x2009;0.007). A high mRNA-score was linked with more "resting"/ "anti-inflammatory" rather than "activated"/ "pro-inflammatory" tumor-infiltrating immune cells and ICI resistance in GIACs (GSE165252, P&#x2009;=&#x2009;0.047; PRJEB25780, P&#x2009;=&#x2009;0.047) and other solid tumors such as urothelial carcinoma and clear cell renal cell carcinoma. Our findings demonstrate the utility of the Notch pathway in predicting prognosis and ICI resistance. Further studies are warranted to explore the efficacy of Notch inhibitors as immunotherapeutic adjuvants to overcome ICI resistance.

Humans↗

The use of early and midpoint adenoma-carcinoma sequence biomarkers in prediction of neoplastic progression in patients with a history of colorectal neoplasia.

Since significant neoplasia after initial colonoscopy is low, we conducted this pilot study to compare the predictive role for colorectal neoplasia recurrence of anti-DCC with that of Adnab-9 binding to colonic effluent of high-risk patients. DCC and Adnab-9 effluent ELISA were performed at baseline colonoscopies. The results of follow-up colonoscopies were reviewed. To ensure specificity, immunohistochemistry and Western blot was performed with anti-DCC and for Adnab-9 where optimal fixation times were also evaluated. Mean follow-up was 2.6 years. Of 21 patients, 6 of 10 who progressed to CRN and 2 of 11 who did not had a positive Adnab-9 ELISA result (P=0.08). Despite an initial good correlation with Adnab-9 ELISA results in a smaller dataset, we were unable to obtain consistent subsequent DCC immunohistochemistry or Western blot data using antibody from two different sources. However, the original dataset of Adnab-9 results was reproducible on repetition of the ELISA with a larger set of samples that included this initial dataset and optimal fixation time was 20 min. We conclude that Adnab-9 appears to be a promising prognostic marker for neoplasia in the high-risk population. Industry standards need to be developed for DCC monoclonal antibodies that may have similar utility.

Adenoma↗

Patterns of Hg bioaccumulation and transfer in aquatic food webs across multi-lake studies in the northeast US.

The northeastern USA receives some of the highest levels of atmospheric mercury deposition of any region in North America. Moreover, fish from many lakes in this region carry Hg burdens that present health risks to both human and wildlife consumers. The overarching goal of this study was to identify the attributes of lakes in this region that are most likely associated with high Hg burdens in fish. To accomplish this, we compared data collected in four separate multi-lake studies. Correlations among Hg in fish (4 studies) or in zooplankton and fish (2 studies) and numerous chemical, physical, land use, and ecological variables were compared across more than 150 lakes. The analysis produced three general findings. First, the most important predictors of Hg burdens in fish were similar among datasets. As found in past studies, key chemical covariates (e.g., pH, acid neutralizing capacity, and SO4) were negatively correlated with Hg bioaccumulation in the biota. However, negative correlations with several parameters that have not been previously identified (e.g., human land use variables and zooplankton density) were also found to be equally important predictors. Second, certain predictors were unique to individual datasets and differences in lake population characteristics, sampling protocols, and fish species in each study likely explained some of the contrasting results that we found in the analyses. Third, lakes with high rates of Hg bioaccumulation and trophic transfer have low pH and low productivity with relatively undisturbed watersheds suggesting that atmospheric deposition of Hg is the dominant or sole source of input. This study highlights several fundamental complexities when comparing datasets over different environmental conditions but also underscores the utility of such comparisons for revealing key drivers of Hg trophic transfer among different types of lakes.

Animals↗

Genetic algorithms and self-organizing maps: a powerful combination for modeling complex QSAR and QSPR problems.

Modeling non-linear descriptor-target activity/property relationships with many dependent descriptors has been a long-standing challenge in the design of biologically active molecules. In an effort to address this problem, we couple the supervised self-organizing map with the genetic algorithm. Although self-organizing maps are non-linear and topology-preserving techniques that hold great potential for modeling and decoding relationships, the large number of descriptors in typical quantitative structure-activity relationship or quantitative structure-property relationship analysis may lead to spurious correlation(s) and/or difficulty in the interpretation of resulting models. To reduce the number of descriptors to a manageable size, we chose the genetic algorithm for descriptor selection because of its flexibility and efficiency in solving complex problems. Feasibility studies were conducted using six different datasets, of moderate-to-large size and moderate-to-great diversity; each with a different biological endpoint. Since favorable training set statistics do not necessarily indicate a highly predictive model, the quality of all models was confirmed by withholding a portion of each dataset for external validation. We also address the variability introduced onto modeling through dataset partitioning and through the stochastic nature of the combined genetic algorithm supervised self-organizing map method using the z-score and other tests. Experiments show that the combined method provides comparable accuracy to the supervised self-organizing map alone, but using significantly fewer descriptors in the models generated. We observed consistently better results than partial least squares models. We conclude that the combination of genetic algorithms with the supervised self-organizing map shows great potential as a quantitative structure-activity/property relationship modeling tool.

Algorithms↗

Probabilistic Identification of Spin Systems and their Assignments including Coil-Helix Inference as Output (PISTACHIO).

We present a novel automated strategy (PISTACHIO) for the probabilistic assignment of backbone and sidechain chemical shifts in proteins. The algorithm uses peak lists derived from various NMR experiments as input and provides as output ranked lists of assignments for all signals recognized in the input data as constituting spin systems. PISTACHIO was evaluated by comparing its performance with raw peak-picked data from 15 proteins ranging from 54 to 300 residues; the results were compared with those achieved by experts analyzing the same datasets by hand. As scored against the best available independent assignments for these proteins, the first-ranked PISTACHIO assignments were 80-100% correct for backbone signals and 75-95% correct for sidechain signals. The independent assignments benefited, in a number of cases, from structural data (e.g. from NOESY spectra) that were unavailable to PISTACHIO. Any number of datasets in any combination can serve as input. Thus PISTACHIO can be used as datasets are collected to ascertain the current extent of secure assignments, to identify residues with low assignment probability, and to suggest the types of additional data needed to remove ambiguities. The current implementation of PISTACHIO, which is available from a server on the Internet, supports input data from 15 standard double- and triple-resonance experiments. The software can readily accommodate additional types of experiments, including data from selectively labeled samples. The assignment probabilities can be carried forward and refined in subsequent steps leading to a structure. The performance of PISTACHIO showed no direct dependence on protein size, but correlated instead with data quality (completeness and signal-to-noise). PISTACHIO represents one component of a comprehensive probabilistic approach we are developing for the collection and analysis of protein NMR data.

Algorithms↗

3D elastic registration of vessel structures from IVUS data on biplane angiography.

RATIONALE AND OBJECTIVES: Planar angiograms and intravascular ultrasound (IVUS) imaging provide important insight for the evaluation of atherosclerotic diseases and blood flow abnormalities. The construction of realistic three-dimensional models is essential to efficiently follow the progression of arterial plaque. This requires an explicit localization of IVUS frames from angiograms. Because of the difficulties encountered when trying to track the position of an IVUS transducer, we propose an elastic registration approach that relies on a virtual catheter path. MATERIALS AND METHODS: Deformable surface models of the lumen and external wall are constructed from segmented IVUS contours. A crude registration is obtained using a three-dimensional vessel centerline, reconstructed from two calibrated angiograms. Robust optimization of the virtual catheter path, position, absolute orientation, and regulation of the external wall shape is performed until near-perfect alignment of the back-projected silhouettes on image edges is reached. RESULTS: Visual assessment of the reconstructed vessels showed a good superposition of virtual models on the angiograms. We measured a 0.4-mm residual error value. A preliminary study of convergence properties on 15 datasets showed that initial absolute orientation may affect the solution. However, for follow-ups, coherent solutions were found among datasets. CONCLUSION: The advantages of the virtual catheter path approach are demonstrated. Future work will look at ways to single out the true solution with a better use of the available information in both modalities and additional validation studies on improved datasets.

Algorithms↗

Quantitative analysis of brain asymmetry by using the divergence measure: normal-pathological brain discrimination.

RATIONALE AND OBJECTIVES: The human brain demonstrates approximate bilateral symmetry of anatomy, function, neurochemical activity, and electrophysiology. This symmetry reflected in radiological images may be affected by pathology. Hence quantitative analysis of brain symmetry may enable the normal and pathological brain discrimination. We propose a method based on the Jeffreys divergence measure (J-divergence), which attempts to quantify "approximate symmetry" and also aids to classify the brain as bilaterally symmetrical/asymmetrical (normal/abnormal). MATERIALS AND METHODS: The dataset included studies of 101 patients (59 without detectable pathologies and 42 with different abnormalities). First, the midsagittal plane is computed for the volume data that divides the head into two hemispheres. The J-divergence is calculated from the density functions of intensities of both the hemispheres. Statistical analysis was conducted to find the best distribution for normal/abnormal datasets. RESULTS: Statistical tests showed that the lognormal distribution best characterizes the values of the J-divergence for both normal and abnormal cases, and the threshold value for the Jeffreys divergence measure to classify the brains with and without detectable pathologies is T = 0.007. The threshold value had a sensitivity of 88.1% and specificity of 90.9%. CONCLUSION: The proposed method is fast and simple to compute. The high sensitivity and specificity indicate the results are encouraging. This method can be used for the initial analysis of data, detection of pathology, classification of dataset as presumably normal/abnormal, and localization of abnormality.

Algorithms↗

Lung deformation estimation and four-dimensional CT lung reconstruction.

RATIONALE AND OBJECTIVES: Four-dimensional (4D) computed tomography (CT) can be used in radiation treatment planning to account for respiratory motion. Current 4D CT techniques have limitations in either spatial or temporal resolution. In addition, most of these techniques rely on auxiliary surrogates to relate the time of the CT scan to the patient's respiratory phase. We propose a 4D CT method for lung applications to overcome these problems. MATERIALS AND METHODS: A set of axial scans are taken at multiple table positions to obtain a series of two-dimensional images while the patient is breathing freely. Each two-dimensional image is registered to a reference CT volume. The deformation of the image with respect to the volume is used to synchronize the image with the respiratory cycle assuming that there is no phase variation along the craniocaudal direction. The reconstructed 4D dataset is a series of deformable transformations of the reference volume. RESULTS: A synthetic 4D dataset showed that the registration error is less than 5% of the image deformation. A swine study showed that the algorithm can generate better image quality than the image sorting method. A respiratory-gated 4D dataset showed that the algorithm's result is consistent with the ground truth. CONCLUSION: The algorithm can reconstruct good quality 4D images without external surrogates even if the CT scans are acquired under irregular respiratory motion. The algorithm may allow for reduced radiation dose to the patient with a limited loss of image quality. Although the phase variation exists along the craniocaudal direction, the 4D reconstruction is reasonably accurate.

Algorithms↗

Filter versus wrapper gene selection approaches in DNA microarray domains.

DNA microarray experiments generating thousands of gene expression measurements, are used to collect information from tissue and cell samples regarding gene expression differences that could be useful for diagnosis disease, distinction of the specific tumor type, etc. One important application of gene expression microarray data is the classification of samples into known categories. As DNA microarray technology measures the gene expression en masse, this has resulted in data with the number of features (genes) far exceeding the number of samples. As the predictive accuracy of supervised classifiers that try to discriminate between the classes of the problem decays with the existence of irrelevant and redundant features, the necessity of a dimensionality reduction process is essential. We propose the application of a gene selection process, which also enables the biology researcher to focus on promising gene candidates that actively contribute to classification in these large scale microarrays. Two basic approaches for feature selection appear in machine learning and pattern recognition literature: the filter and wrapper techniques. Filter procedures are used in most of the works in the area of DNA microarrays. In this work, a comparison between a group of different filter metrics and a wrapper sequential search procedure is carried out. The comparison is performed in two well-known DNA microarray datasets by the use of four classic supervised classifiers. The study is carried out over the original-continuous and three-intervals discretized gene expression data. While two well-known filter metrics are proposed for continuous data, four classic filter measures are used over discretized data. The same wrapper approach is used for both continuous and discretized data. The application of filter and wrapper gene selection procedures leads to considerably better accuracy results in comparison to the non-gene selection approach, coupled with interesting and notable dimensionality reductions. Although the wrapper approach mainly shows a more accurate behavior than filter metrics, this improvement is coupled with considerable computer-load necessities. We note that most of the genes selected by proposed filter and wrapper procedures in discrete and continuous microarray data appear in the lists of relevant-informative genes detected by previous studies over these datasets. The aim of this work is to make contributions in the field of the gene selection task in DNA microarray datasets. By an extensive comparison with more popular filter techniques, we would like to make contributions in the expansion and study of the wrapper approach in this type of domains.

Artificial Intelligence↗

Applying spatial distribution analysis techniques to classification of 3D medical images.

OBJECTIVE: The objective of this paper is to classify 3D medical images by analyzing spatial distributions to model and characterize the arrangement of the regions of interest (ROIs) in 3D space. METHODS AND MATERIAL: Two methods are proposed for facilitating such classification. The first method uses measures of similarity, such as the Mahalanobis distance and the Kullback-Leibler (KL) divergence, to compute the difference between spatial probability distributions of ROIs in an image of a new subject and each of the considered classes represented by historical data (e.g., normal versus disease class). A new subject is predicted to belong to the class corresponding to the most similar dataset. The second method employs the maximum likelihood (ML) principle to predict the class that most likely produced the dataset of the new subject. RESULTS: The proposed methods have been experimentally evaluated on three datasets: synthetic data (mixtures of Gaussian distributions), realistic lesion-deficit data (generated by a simulator conforming to a clinical study), and functional MRI activation data obtained from a study designed to explore neuroanatomical correlates of semantic processing in Alzheimer's disease (AD). CONCLUSION: Performed experiments demonstrated that the approaches based on the KL divergence and the ML method provide superior accuracy compared to the Mahalanobis distance. The later technique could still be a method of choice when the distributions differ significantly, since it is faster and less complex. The obtained classification accuracy with errors smaller than 1% supports that useful diagnosis assistance could be achieved assuming sufficiently informative historic data and sufficient information on the new subject.

Algorithms↗

Improved and automated prediction of effective siRNA.

Short interfering RNAs are used in functional genomics studies to knockdown a single gene in a reversible manner. The results of siRNA experiments are highly dependent on the choice of siRNA sequence. In order to evaluate siRNA design rules, we collected a database of 398 siRNAs of known efficacy from 92 genes. We used this database to evaluate previously proposed rules from smaller datasets, and to find a new set of rules that are optimal for the entire database. We also trained a regression tree with full cross-validation. It was however difficult to obtain the same precision as methods previously tested on small datasets from one or two genes. We show that those methods are overfitting as they work poorly on independent validation datasets from multiple genes. Our new design rules can predict siRNAs with efficacy >/= 50% in 91% of cases, and with efficacy >/=90% in 52% of cases, which is more than a twofold improvement over random selection. Software for designing siRNAs is available online via a web server at or as a standalone version for high-throughput applications.

Algorithms↗

Prediction of protein subcellular locations by GO-FunD-PseAA predictor.

The localization of a protein in a cell is closely correlated with its biological function. With the explosion of protein sequences entering into DataBanks, it is highly desired to develop an automated method that can fast identify their subcellular location. This will expedite the annotation process, providing timely useful information for both basic research and industrial application. In view of this, a powerful predictor has been developed by hybridizing the gene ontology approach [Nat. Genet. 25 (2000) 25], functional domain composition approach [J. Biol. Chem. 277 (2002) 45765], and the pseudo-amino acid composition approach [Proteins Struct. Funct. Genet. 43 (2001) 246; Erratum: ibid. 44 (2001) 60]. As a showcase, the recently constructed dataset [Bioinformatics 19 (2003) 1656] was used for demonstration. The dataset contains 7589 proteins classified into 12 subcellular locations: chloroplast, cytoplasmic, cytoskeleton, endoplasmic reticulum, extracellular, Golgi apparatus, lysosomal, mitochondrial, nuclear, peroxisomal, plasma membrane, and vacuolar. The overall success rate of prediction obtained by the jackknife cross-validation was 92%. This is so far the highest success rate performed on this dataset by following an objective and rigorous cross-validation procedure.

Algorithms↗

Quantitative structure-activity relationships for small non-peptide antagonists of CXCR2: indirect 3D approach using the frontal polygon method.

The chemokine receptor, CXCR2, plays an important role in recruiting granulocytes to sites of inflammation and has been proposed as an important therapeutic target. A number of CXCR2 antagonists have been synthesized and evaluated; however, quantitative structure-activity relationship (QSAR) models have not been developed for these molecules. Most CXCR2 antagonists can be grouped into four related categories: N,N'-diphenylureas, nicotinamide N-oxides, quinoxalines, and triazolethiols. Based on these categories, we developed a QSAR model for 59 nonpeptide antagonists of CXCR2 using a partial 3D comparison of the antagonists with local fingerprints obtained from rigid and flexible fragments of the molecules. Each compound was represented by calculated structural descriptors that encoded atomic charge, molar refraction, hydrophobicity, and geometric features. We obtained good conventional R(2) coefficients, high leave-one-out cross-validated values for the whole dataset (R(cv)(2)=0.785), as well as for the dataset divided into subsets of triazolethiol derivatives (R(cv)(2)=0.821) and joint subset of N'-diphenylureas, nicotinamide N-oxides, N,N'-diphenylureas, and quinoxaline derivatives and quinoxalines derivatives (R(cv)(2)=0.766), indicating a good predictive ability and robustness of the model. Additionally, charge distribution was found to be a significant contributor in modeling whole dataset. Using our model, structural fragments (submolecules) responsible for the antagonist activity were also identified. These data suggest the QSAR models developed here may be useful in guiding the design of CXCR2 antagonists from molecular fragments.

Granulocytes↗

Constructing plasma protein binding model based on a combination of cluster analysis and 4D-fingerprint molecular similarity analyses.

Based on 2D-connectivity molecular similarity and cluster analyses, a dataset for HSA binding is divided into the training set and the test set. 4D-fingerprint similarity measures were applied to this dataset. Four different predictive schemes (SM, SA, SR, and SC) were applied to the test set based on the similarity measures of each compound to the compounds in the training set. The first algorithmic scheme (SM), which only takes the most similar compound in the training set into consideration, predicts the binding affinity of a test compound. This scheme has relatively poor predictivity based on 4D-fingerprint similarity analyses. The other three algorithmic schemes (SM, SR, and SC), which assign a weighting coefficient to each of the top-ten most similar training set compounds, have reasonable predictivity of a test set. The algorithmic scheme which categorizes the most similar compounds into different weighted clusters predicts the test set best. The 4D-fingerprints provide 36 different individual IPE/IPE type molecular similarity measures. Further investigation shows that the NP/HA, HS/HA, and HA/HA IPE/IPE type measures predict the test set well. Moreover, these three IPE/IPE type similarity measures are very similar to one another for the particular training and test sets investigated. The 4D-fingerprints have relatively high predictivity for this particular dataset.

Algorithms↗

Linking the human cytogenetic map with nucleotide sequence: the CCAP clone set.

We present the completed dataset and clone repository of the Cancer Chromosome Aberration Project (CCAP), an initiative developed and funded through the intramural program of the U.S. National Cancer Institute, to provide seamless linkage of human cytogenetic markers with the primary nucleotide sequence of the human genome. Spaced at 1-2 Mb intervals across the human genome, 1,339 bacterial artificial chromosome (BAC) clones have been localized to chromosomal bands through high-resolution fluorescence in situ hybridization (FISH) mapping. Of these clones, 99.8% can be positioned on the primary human genome sequence and 95% are placed at or close to their precise nucleotide starts and stops. This dataset can be studied and manipulated within generally available public Web sites. The clones are available from a commercial repository. The CCAP BAC clone set provides anchors for the interrogation of gene and sequence involvement in oncogenic and developmental disorders when the starting point is the recognition of a structural, numerical, or interstitial chromosomal aberration. This dataset also provides a current view of the quality and coherence of the available genome sequence and insight into the nucleotide and three-dimensional structures that manifest as Giemsa light and dark chromosomal banding patterns.

Base Composition↗

Age-associated variation in lipid composition and nutritional quality of the invasive bivalve Anadara kagoshimensis from the Sea of Azov.

This study examined age-associated variation in lipid composition and nutritional quality in the invasive clam Anadara kagoshimensis from the Sea of Azov. Fatty acid and sterol compositions were determined by gas chromatography in clams assigned by sclerochronology to five age classes (2-6&#xa0;years). Several lipid-quality indices, including PUFA/SFA, IA, IT, hH, and HPI, exhibited their most favorable mean values in the 4-year group. The highest condition index (13.7) and the lowest n-3/n-6 ratio (1.1) were observed in the small and unbalanced group of 6-year-old clams. An exploratory linear discriminant analysis (LDA) workflow was used to examine whether primary lipid variables and mathematically derived composite features could enhance age-group separation. LDA based on primary variables yielded 78% leave-one-out classification within the present dataset. Data-driven composite-feature construction increased the within-dataset classification to 100% and graphical group separation to 85%, the values with yet restricted predictive accuracy due to the limited dataset. The results provide baseline information on age-associated lipid variation in A. kagoshimensis and illustrate a proof-of-concept workflow that can be tested in future studies using larger, balanced samples and special validation approaches.

Arcidae↗