Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Weighted-support vector machines for predicting membrane protein types based on pseudo-amino acid composition.

Membrane proteins are generally classified into the following five types: (1) type I membrane proteins, (2) type II membrane proteins, (3) multipass transmembrane proteins, (4) lipid chain-anchored membrane proteins and (5) GPI-anchored membrane proteins. Prediction of membrane protein types has become one of the growing hot topics in bioinformatics. Currently, we are facing two critical challenges in this area: first, how to take into account the extremely complicated sequence-order effects, and second, how to deal with the highly uneven sizes of the subsets in a training dataset. In this paper, stimulated by the concept of using the pseudo-amino acid composition to incorporate the sequence-order effects, the spectral analysis technique is introduced to represent the statistical sample of a protein. Based on such a framework, the weighted support vector machine (SVM) algorithm is applied. The new approach has remarkable power in dealing with the bias caused by the situation when one subset in the training dataset contains many more samples than the other. The new method is particularly useful when our focus is aimed at proteins belonging to small subsets. The results obtained by the self-consistency test, jackknife test and independent dataset test are encouraging, indicating that the current approach may serve as a powerful complementary tool to other existing methods for predicting the types of membrane proteins.

Algorithms↗

Transmembrane helix prediction: a comparative evaluation and analysis.

The prediction of transmembrane (TM) helices plays an important role in the study of membrane proteins, given the relatively small number (approximately 0.5% of the PDB) of high-resolution structures for such proteins. We used two datasets (one redundant and one non-redundant) of high-resolution structures of membrane proteins to evaluate and analyse TM helix prediction. The redundant (non-redundant) dataset contains structure of 434 (268) TM helices, from 112 (73) polypeptide chains. Of the 434 helices in the dataset, 20 may be classified as 'half-TM' as they are too short to span a lipid bilayer. We compared 13 TM helix prediction methods, evaluating each method using per segment, per residue and termini scores. Four methods consistently performed well: SPLIT4, TMHMM2, HMMTOP2 and TMAP. However, even the best methods were in error by, on average, about two turns of helix at the TM helix termini. The best and worst case predictions for individual proteins were analysed. In particular, the performance of the various methods and of a consensus prediction method, were compared for a number of proteins (e.g. SecY, ClC, KvAP) containing half-TM helices. The difficulties of predicting half-TM helices suggests that current prediction methods successfully embody the two-state model of membrane protein folding, but do not accommodate a third stage in which, e.g., short helices and re-entrant loops fold within a bundle of stable TM helices.

Amino Acid Sequence↗

The infant index: a new outcome measure for pre-school children's services.

BACKGROUND: The evaluation of community services for preschool children is hampered by the lack of valid and routinely available outcome measures. This study examines the use of data collected by teachers in response to educational legislation to determine whether a routine measure of attainments in primary school is sensitive to factors known to affect mental development. METHOD: A community child health dataset for the cohort of children born in Sheffield in 1990-1991 was matched with a dataset provided by schools in 1995-1996. The educational data consisted of the Infant Index scores which measure education attainments in reception class pupils. RESULTS: We matched 4487 children from both datasets, which represented 75 per cent of all children born in the 1990-1991 cohort. Factors which predicted a poor Infant Index included male gender (odds ratio (OR) = 2.1, 95 per cent confidence interval (CI)= 1.8-2.6), low birthweight (OR = 1.4, 95 per cent CI = 1.1-1.9) and lack of breast feeding either by intention to feed (OR = 1.3, 95 per cent CI = 1.1-1.7) or actual feeding practice at one month (OR = 1.5, 95 per cent CI = 1.1-2.0). Other factors associated with a poor outcome for the child were postnatal depression, number of pregnancies, ethnicity, pre-school educational experiences and poor housing. CONCLUSIONS: Although the results are interesting in themselves, the main significance of our project is in establishing a link between routinely collected health data and routine education data. This could facilitate research in the future thus leading to a considerable saving in the cost of long-term intervention studies.

Birth Weight↗

The Hunter Serotonin Toxicity Criteria: simple and accurate diagnostic decision rules for serotonin toxicity.

BACKGROUND: There are difficulties with the diagnosis of serotonin toxicity, particularly with the use of Sternbach's criteria. AIM: To improve the criteria for diagnosing clinically significant serotonin toxicity. DESIGN: Retrospective analysis of prospectively collected data METHODS: We studied all patients admitted to the Hunter Area Toxicology Service (HATS) following an overdose of a serotonergic drug from January 1987 to November 2002 (n = 2222). Main outcomes were: diagnosis of serotonin toxicity by a clinical toxicologist, fulfillment of Sternbach's criteria and treatment with a serotonin receptor (5-HT(2A)) antagonist. A learning dataset of 473 selective serotonin reuptake inhibitor (SSRI)-alone overdoses was used to determine individual clinical features predictive of serotonin toxicity by univariate analysis. Decision rules using CART analysis were developed, and tested on the dataset of all serotonergic overdose admissions. RESULTS: Numerous clinical features were associated with serotonin toxicity, but only clonus (inducible, spontaneous or ocular), agitation, diaphoresis, tremor and hyperreflexia were needed for accurate prediction of serotonin toxicity as diagnosed by a clinical toxicologist. Although the learning dataset did not include patients with life-threatening serotonin toxicity, hypertonicity and maximum temperature > 38 degrees C were universal in such patients; these features were therefore added. Using these seven clinical features, decision rules (the Hunter Serotonin Toxicity Criteria) were developed. These new criteria were simpler, more sensitive (84% vs. 75%) and more specific (97% vs. 96%) than Sternbach's criteria. DISCUSSION: These redefined criteria for serotonin toxicity should be more sensitive to serotonin toxicity and less likely to yield false positives.

Analysis of Variance↗

Environmental benzene exposure induces a conserved neutrophil degranulation program across species.

Immune systems have evolved under constant pressure from pathogens and environmental challenges, leading to the emergence of conserved defense mechanisms across diverse organisms. Evidence indicates that environmental exposures perturb immune regulatory networks, particularly during development, when transcriptional programs governing hematopoiesis, immune cell differentiation, and inflammatory signaling are highly dynamic and sensitive to external stressors. Volatile organic compounds represent an important but incompletely understood source of immunological perturbation. Among these, benzene is a ubiquitous environmental contaminant associated with hematotoxicity and immune dysregulation; however, transcriptional responses to environmentally relevant low-level exposures during development remain poorly characterized. To determine whether benzene exposure engages conserved cross-species immune regulatory pathways, we performed a comparative transcriptomic analysis integrating developmental tissues from 3 vertebrate systems: human placenta, murine placenta, and zebrafish larvae. Bulk RNA sequencing datasets were analyzed to identify transcriptional responses associated with benzene exposure in experimental models (≤5 ppm) and with benzene adduct levels in maternal plasma for human samples. Because placental gene expression exhibits strong sexual dimorphism, murine datasets were stratified by fetal sex. Pathway- and network-level analyses were used to identify conserved biological responses. We observed a striking convergence on activation of innate immune pathways associated with neutrophil degranulation, IL-8 signaling, and Rho GTPase-mediated inflammatory responses. Further, network analyses identified CXCL8 and ERK1/2 as shared regulatory hubs linking transcriptional responses across datasets. Together, these findings uncover an evolutionarily conserved innate immune signature associated with benzene exposure during vertebrate development, suggesting that environmental chemical perturbations may disrupt fundamental immune regulatory programs across species.

Animals↗

Registration of three-dimensional magnetic resonance and radionuclide skeletal images.

PURPOSE: To apply postprocessing techniques to register three-dimensional TI-201 bone SPECT datasets with MRI. This may provide more accurate anatomic-functional correlation when localizing active tumors. MATERIALS AND METHODS: Three-dimensional datasets were constructed from previously acquired MRIs using routine imaging protocols. Registration software was used to coregister the TI-201 SPECT studies and the MRIs in three dimensions. RESULTS: Adequate TI-201 uptake in muscles and soft tissues along with relatively low accumulation in tendons and joint spaces provided adequate landmarks for visually aligning SPECT and MRI datasets. MR abnormalities were more extensive because of surrounding reactive tissue, and more focal TI-201 uptake could be demonstrated within the region of MR signal abnormality, allowing the focal metabolically active tissue to be distinguished from adjacent edema. CONCLUSIONS: Image registration of SPECT and anatomic imaging (CT or MRI) is used routinely to evaluate functional abnormalities within the brain. This technique has now been applied to the combination of TI-201 SPECT and MR data for evaluating bone lesions and may provide additional anatomic information for localizing functional abnormalities. This may be valuable for defining targets for biopsy, planning surgical treatment, and using minimally invasive therapies.

Bone Neoplasms↗

Linked insurance-tumor registry database for health services research.

OBJECTIVE: Breast cancer screening and treatment data are often limited to restricted populations, including women older than 65 years old. The goal of this project was to develop procedures to link tumor registry and insurance claims databases on women younger than 65 years old with breast cancer and to assess the accuracy and validity of the linked dataset. METHODS: Iowa Cancer Registry (ICR) and Wellmark Blue Cross/Blue Shield of Iowa (BC/BS) membership files of women with incident in situ or invasive breast cancer from 1989 to 1996 were linked. An automated deterministic match was followed with visual inspection from three independent reviewers applying a matching protocol. Matched and overall registry data were compared to assess population representativeness. Claims from BC/BS for incident cases during 1994 were examined for coding of a recent breast cancer diagnosis or treatment. RESULTS: The final dataset included 4,397 matched cases of patients aged 21 years and older from 1989 to 1996. The sociodemographic and tumor characteristics of the ICR population younger than 65 years old (n = 7,469) with breast cancer or carcinoma in situ were nearly identical with those of the matched patients younger than 65 years old (n = 3,449). Nearly all (96%) of the 445 matched incident cases in 1994 had claims data (CPT, DRG, or ICD-9 code) indicative of breast cancer. Treatment patterns varied by data source, with agreement ranging from 76% to 82%. CONCLUSIONS: The validity and generalizability of these data demonstrate their potential for further health services research among younger insured women with breast cancer. Additionally, the process outlined may be useful for developing other datasets to study other cancers in the population younger than 65 years old.

Adult↗

Cumulative meta-analysis: a new tool for detection of temporal trends and publication bias in ecology.

Temporal changes in the magnitude of research findings have recently been recognized as a general phenomenon in ecology, and have been attributed to the delayed publication of non-significant results and disconfirming evidence. Here we introduce a method of cumulative meta-analysis which allows detection of both temporal trends and publication bias in the ecological literature. To illustrate the application of the method, we used two datasets from recently conducted meta-analyses of studies testing two plant defence theories. Our results revealed three phases in the evolution of the treatment effects. Early studies strongly supported the hypothesis tested, but the magnitude of the effect decreased considerably in later studies. In the latest studies, a trend towards an increase in effect size was observed. In one of the datasets, a cumulative meta-analysis revealed publication bias against studies reporting disconfirming evidence; such studies were published in journals with a lower impact factor compared to studies with results supporting the hypothesis tested. Correlation analysis revealed neither temporal trends nor evidence of publication bias in the datasets analysed. We thus suggest that cumulative meta-analysis should be used as a visual aid to detect temporal trends and publication bias in research findings in ecology in addition to the correlative approach.

Carbon↗

Does a tree-like phylogeny only exist at the tips in the prokaryotes?

The extent to which prokaryotic evolution has been influenced by horizontal gene transfer (HGT) and therefore might be more of a network than a tree is unclear. Here we use supertree methods to ask whether a definitive prokaryotic phylogenetic tree exists and whether it can be confidently inferred using orthologous genes. We analysed an 11-taxon dataset spanning the deepest divisions of prokaryotic relationships, a 10-taxon dataset spanning the relatively recent gamma-proteobacteria and a 61-taxon dataset spanning both, using species for which complete genomes are available. Congruence among gene trees spanning deep relationships is not better than random. By contrast, a strong, almost perfect phylogenetic signal exists in gamma-proteobacterial genes. Deep-level prokaryotic relationships are difficult to infer because of signal erosion, systematic bias, hidden paralogy and/or HGT. Our results do not preclude levels of HGT that would be inconsistent with the notion of a prokaryotic phylogeny. This approach will help decide the extent to which we can say that there is a prokaryotic phylogeny and where in the phylogeny a cohesive genomic signal exists.

Bacteria↗

'The surface management system' (SuMS) database: a surface-based database to aid cortical surface reconstruction, visualization and analysis.

Surface reconstructions of the cerebral cortex are increasingly widely used in the analysis and visualization of cortical structure, function and connectivity. From a neuroinformatics perspective, dealing with surface-related data poses a number of challenges. These include the multiplicity of configurations in which surfaces are routinely viewed (e.g. inflated maps, spheres and flat maps), plus the diversity of experimental data that can be represented on any given surface. To address these challenges, we have developed a surface management system (SuMS) that allows automated storage and retrieval of complex surface-related datasets. SuMS provides a systematic framework for the classification, storage and retrieval of many types of surface-related data and associated volume data. Within this classification framework, it serves as a version-control system capable of handling large numbers of surface and volume datasets. With built-in database management system support, SuMS provides rapid search and retrieval capabilities across all the datasets, while also incorporating multiple security levels to regulate access. SuMS is implemented in Java and can be accessed via a Web interface (WebSuMS) or using downloaded client software. Thus, SuMS is well positioned to act as a multiplatform, multi-user 'surface request broker' for the neuroscience community.

Brain Mapping↗

Genomics and plant cells: application of genomics strategies to Arabidopsis cell biology.

In this review I seek to describe how the complete catalogue of plant genes and proteins, revealed by genome sequencing, can provide novel insights into cell biology. Many new analytical methods have been developed to digest the flood of genome sequence data, including analysis of the transcriptome, proteome and metabolites. High-throughput analysis of protein targeting and other methods will ascribe new information to proteins and create important links with other large datasets. To fulfil the potential revealed by this genomic information, many challenges have to be met. Among these are organizational changes needed to create common datasets accessible to all scientists, and bioinformatics solutions to capture and integrate diverse datasets. Once harnessed, these new strategies will irrevocably change the way we conduct plant science.

Arabidopsis↗

TaxI: a software tool for DNA barcoding using distance methods.

DNA barcoding is a promising approach to the diagnosis of biological diversity in which DNA sequences serve as the primary key for information retrieval. Most existing software for evolutionary analysis of DNA sequences was designed for phylogenetic analyses and, hence, those algorithms do not offer appropriate solutions for the rapid, but precise analyses needed for DNA barcoding, and are also unable to process the often large comparative datasets. We developed a flexible software tool for DNA taxonomy, named TaxI. This program calculates sequence divergences between a query sequence (taxon to be barcoded) and each sequence of a dataset of reference sequences defined by the user. Because the analysis is based on separate pairwise alignments this software is also able to work with sequences characterized by multiple insertions and deletions that are difficult to align in large sequence sets (i.e. thousands of sequences) by multiple alignment algorithms because of computational restrictions. Here, we demonstrate the utility of this approach with two datasets of fish larvae and juveniles from Lake Constance and juvenile land snails under different models of sequence evolution. Sets of ribosomal 16S rRNA sequences, characterized by multiple indels, performed as good as or better than cox1 sequence sets in assigning sequences to species, demonstrating the suitability of rRNA genes for DNA barcoding.

Animals↗

Genomic characterization and phylogenetic placement of Matryoshka RNA virus 1 associated with Plasmodium vivax malaria in Africa.

Plasmodium vivax is a major cause of human malaria. It harbours Matryoshka RNA virus 1 (MaRNAV-1), a bi-segmented positive-sense RNA virus. MaRNAV-1 was first described in P. vivax and is now recognized as part of a wider group of Matryoshka viruses. These viruses also infect other haemosporidian parasites such as Leucocytozoon and Haemoproteus. The presence of MaRNAV-1 in African-origin human P. vivax, however, has not been clearly established. This study investigated whether MaRNAV-1 is present in public African-origin P. vivax transcriptomic datasets. Any viral sequences recovered were characterized using comparative genomic and phylogenetic analyses. A secondary in silico analysis targeted African-origin P. vivax RNA-seq runs from public repositories. Although the search covered Africa, only Ethiopian datasets could be confidently identified, retrieved and compiled at the time. After quality control and screening for MaRNAV-1 RNA-dependent RNA polymerase (RdRp) signals, three high-confidence runs were selected for further analysis. Reference-guided reconstruction, ORF prediction, blast-based validation and RdRp phylogenetic analysis were performed. MaRNAV-1 was identified in three Ethiopian P. vivax malaria transcriptomes. This was supported by strong segment-level mapping, near-complete coverage, high mean depth and minimal low-depth masking. The recovered genomes showed the expected bisegmented organization of MaRNAV-1. Segment I was highly conserved and encoded the canonical RdRp in all three consensus sequences. Segment II showed the conserved organization of two overlapping hypothetical ORFs in all three consensus sequences. Blast analyses confirmed close similarity to MaRNAV-1 reference sequences. Phylogenetic inference grouped the Ethiopian sequences within the broader P. vivax-associated MaRNAV-1 lineage, alongside other recognized MaRNAV lineages distinct from more divergent narna-like viruses. These findings provide genomic evidence for MaRNAV-1 in publicly available African-origin P. vivax transcriptomic datasets and add to the emerging evidence for the virus in the African malaria context.

MaRNAV↗

Low linkage disequilibrium indicative of recombination in foot-and-mouth disease virus gene sequence alignments.

We have applied tests for detecting recombination to genes of foot-and-mouth disease virus (FMDV). Our approach estimated summary statistics of linkage disequilibrium (LD), which are sensitive to recombination. Using the genealogical relationships, rate heterogeneity and mutation parameters estimated from individual sets of aligned gene sequences, we simulated matching RNA sequence datasets without recombination. These simulated datasets allowed for recurrent mutations at any site to mimic homoplasy in virus sequence data and allow construction of null distributions for LD parameters expected in the absence of recombination. We tested for recombination in two ways: by comparing LD in observed data with corresponding null distributions obtained from simulated data; and by testing for a negative relationship between observed LD between pairs of polymorphic nucleotide sites and inter-site distance. We applied these tests to six FMDV datasets from four serotypes and found some evidence for recombination in all of them.

Capsid Proteins↗

Rhinovirus infection of airway epithelial cells uncovers the non-ciliated subset as a likely driver of genetic susceptibility to childhood-onset asthma.

Asthma is a complex disease caused by genetic and environmental factors. Epidemiological studies have shown that in children, wheezing during rhinovirus infection (a cause of the common cold) is associated with asthma development during childhood. This has led scientists to hypothesize there could be a causal relationship between rhinovirus infection and asthma or that RV-induced wheezing identifies individuals at increased risk for asthma development. However, not all children who wheeze when they have a cold develop asthma. Genome-wide association studies (GWAS) have identified hundreds of genetic variants contributing to asthma susceptibility, with the vast majority of likely causal variants being non-coding. Integrative analyses with transcriptomic and epigenomic datasets have indicated that T cells drive asthma risk, which has been supported by mouse studies. However, the datasets ascertained in these integrative analyses lack airway epithelial cells. Furthermore, large-scale transcriptomic T cell studies have not identified the regulatory effects of most non-coding risk variants in asthma GWAS, indicating there could be additional cell types harboring these "missing regulatory effects". Given that airway epithelial cells are the first line of defense against rhinovirus, we hypothesized they could be mediators of genetic susceptibility to asthma. Here we integrate GWAS data with transcriptomic datasets of airway epithelial cells subject to stimuli that could induce activation states relevant to asthma. We demonstrate that epithelial cultures infected with rhinovirus significantly upregulate childhood-onset asthma-associated genes. We show that this upregulation occurs specifically in non-ciliated epithelial cells. This enrichment for genes in asthma risk loci, or 'asthma heritability enrichment' is also significant for epithelial genes upregulated with influenza infection, but not with SARS-CoV-2 infection or cytokine activation. Additionally, cells from patients with asthma showed a stronger heritability enrichment compared to cells from healthy individuals. Overall, our results suggest that rhinovirus infection is an environmental factor that interacts with genetic risk factors through non-ciliated airway epithelial cells to drive childhood-onset asthma.

Preprint↗

Genome-wide association study of copy number variations in Parkinson's disease.

OBJECTIVE: To investigate the impact of copy number variations (CNVs) on Parkinson's disease (PD) pathogenesis using genome-wide data and explore their role in sporadic PD. METHODS: We analyzed CNV data from 11,035 PD patients (including 2,731 early-onset PD (EOPD)) and 8,901 controls from the COURAGE-PD consortium using a sliding window CNV-GWAS and genome-wide burden analysis. The independent dataset from the Global Parkinson Genetics Program (GP2) consisted of 23,089 cases and 18,824 controls were used to validate our initial findings. RESULTS: The exploratory dataset identifies multiple CNV regions associated with PD risk. The nominated CNV loci were not confirmed in an independent dataset, except that only a deletion in the PRKN gene, a well-established EOPD locus, remained genome-wide significant and robustly supported. CNV burden analysis showed a higher prevalence of CNVs in PD-related genes in patients compared to controls (OR=1.56 [1.18-2.09], p=0.0013), with PRKN showing the highest burden (OR=1.47 [1.10-1.98], p=0.026). Patients with CNVs in PRKN had an earlier disease onset. Burden analysis with controls and EOPD patients showed similar results. INTERPRETATION: The largest CNV-based GWAS on PD highlights both the promise and pitfalls of array-based CNV detection in PD and underscores the relevance of whole-genome sequencing approaches in resolving the role of CNV in PD. The array-based findings are prone towards false positive findings that might arise either from platform limitations and/or cohort biases. Future studies require improved genotyping resolution and rigorous cross-cohort validation to reliably assess CNV contributions to PD risk.

Journal Article↗

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Journal Article↗

An Updated Polygenic Index Repository: Expanded Phenotypes, New Cohorts, and Improved Causal Inference.

Polygenic indexes (PGIs) - DNA-based predictors of individual phenotypes - have become essential tools across biomedical and social sciences. We introduce Version 2 of the Polygenic Index Repository, which expands phenotype coverage from 47 to 61, increases the number of participating datasets from 11 to 20, and adopts a more consistent and improved methodology for PGI construction. For 16 phenotypes, we leverage summary statistics from an updated GWAS meta-analysis with greater statistical power compared to the original release, thereby improving the PGI's predictive power. To improve power for family-based analyses, we provide imputed parental PGIs in all datasets with first-degree relatives and offer a framework for interpreting results from analyses that control for parental PGIs. We illustrate the utility of parental PGIs using two applications: (1) comparing PGI associations with and without parental PGI controls for all phenotypes in two Repository datasets with family data, and (2) for BMI and diastolic blood pressure, exploring the contribution of causal versus non-causal components of PGI associations to the imperfect portability of PGIs across subgroups within a genetic ancestry. Collectively, the updates enhance predictive performance, broaden the Repository's scope, and introduce novel resources that reduce confounding bias and improve interpretability.

Journal Article↗