Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

A combined analysis of D22S278 marker alleles in affected sib-pairs: support for a susceptibility locus for schizophrenia at chromosome 22q12. Schizophrenia Collaborative Linkage Group (Chromosome 22).

Several groups have reported weak evidence for linkage between schizophrenia and genetic markers located on chromosome 22q using the lod score method of analysis. However these findings involved different genetic markers and methods of analysis, and so were not directly comparable. To resolve this issue we have performed a combined analysis of genotypic data from the marker D22S278 in multiply affected schizophrenic families derived from 11 independent research groups worldwide. This marker was chosen because it showed maximum evidence for linkage in three independent datasets (Vallada et al., Am J Med Genet 60:139-146, 1995; Polymeropoulos et al., Neuropsychiatr Genet 54:93-99, 1994; Lasseter et al., Am J Med Genet, 60:172-173, 1995. Using the affected sib-pair method as implemented by the program ESPA, the combined dataset showed 252 alleles shared compared with 188 alleles not share (chi-square 9.31, 1df, P = 0.001) where parental genotype data was completely known. When sib-pairs for whom parental data was assigned according to probability were included the number of alleles shared was 514.1 compared with 437.8 not shared (chi-square 6.12, 1df, P = 0.006). Similar results were obtained when a likelihood ratio method for sib-pair analysis was used. These results indicate that may be a susceptibility locus for schizophrenia at 22q12.

Alleles↗

Prediction of secondary structural content of proteins from their amino acid composition alone. I. New analytic vector decomposition methods.

The predictive limits of the amino acid composition for the secondary structural content (percentage of residues in the secondary structural states helix, sheet, and coil) in proteins are assessed quantitatively. For the first time, techniques for prediction of secondary structural content are presented which rely on the amino acid composition as the only information on the query protein. In our first method, the amino acid composition of an unknown protein is represented by the best (in a least square sense) linear combination of the characteristic amino acid compositions of the three secondary structural types computed from a learning set of tertiary structures. The second technique is a generalization of the first one and takes into account also possible compositional couplings between any two sorts of amino acids. Its mathematical formulation results in an eigenvalue/eigenvector problem of the second moment matrix describing the amino acid compositional fluctuations of secondary structural types in various proteins of a learning set. Possible correlations of the principal directions of the eigenspaces with physical properties of the amino acids were also checked. For example, the first two eigenvectors of the helical eigenspace correlate with the size and hydrophobicity of the residue types respectively. As learning and test sets of tertiary structures, we utilized representative, automatically generated subsets of Protein Data Bank (PDB) consisting of non-homologous protein structures at the resolution thresholds < or = 1.8A, < or = 2.0A, < or = 2.5A, and < or = 3.0 A. We show that the consideration of compositional couplings improves prediction accuracy, albeit not dramatically. Whereas in the self-consistency test (learning with the protein to be predicted), a clear decrease of prediction accuracy with worsening resolution is observed, the jackknife test (leave the predicted protein out) yielded best results for the largest dataset (< or = 3.0A, almost no difference to the self-consistency test!), i.e., only this set, with more than 400 proteins, is sufficient for stable computation of the parameters in the prediction function of the second method. The average absolute error in predicting the fraction of helix, sheet, and coil from amino acid composition of the query protein are 13.7, 12.6, and 11.4%, respectively with r.m.s. deviations in the range of 8.6 divided by 11.8% for the 3.0 A dataset in a jackknife test. The absolute precision of the average absolute errors is in the range of 1 divided by 3% as measured for other representative subsets of the PDB. Secondary structural content prediction methods found in the literature have been clustered in accordance with their prediction accuracies. To our surprise, much more complex secondary structure prediction methods utilized for the same purpose of secondary structural content prediction achieve prediction accuracies very similar to those of the present analytic techniques, implying that all the information beyond the amino acid composition is, in fact, mainly utilized for positioning the secondary structural state in the sequence but not for determination of the overall number of residues in a secondary structural type. This result implies that higher prediction accuracies cannot be achieved relying solely on the amino acid composition of an unknown query protein as prediction input. Our prediction program SSCP has been made available as a World Wide Web and E-mail service.

Amino Acids↗

Evaluation and improvement of multiple sequence methods for protein secondary structure prediction.

A new dataset of 396 protein domains is developed and used to evaluate the performance of the protein secondary structure prediction algorithms DSC, PHD, NNSSP, and PREDATOR. The maximum theoretical Q3 accuracy for combination of these methods is shown to be 78%. A simple consensus prediction on the 396 domains, with automatically generated multiple sequence alignments gives an average Q3 prediction accuracy of 72.9%. This is a 1% improvement over PHD, which was the best single method evaluated. Segment Overlap Accuracy (SOV) is 75.4% for the consensus method on the 396-protein set. The secondary structure definition method DSSP defines 8 states, but these are reduced by most authors to 3 for prediction. Application of the different published 8- to 3-state reduction methods shows variation of over 3% on apparent prediction accuracy. This suggests that care should be taken to compare methods by the same reduction method. Two new sequence datasets (CB513 and CB251) are derived which are suitable for cross-validation of secondary structure prediction methods without artifacts due to internal homology. A fully automatic World Wide Web service that predicts protein secondary structure by a combination of methods is available via http://barton.ebi.ac.uk/.

Algorithms↗

Classification of protein sequences by homology modeling and quantitative analysis of electrostatic similarity.

Protein electrostatics plays a key role in ligand binding and protein-protein interactions. Therefore, similarities or dissimilarities in electrostatic potentials can be used as indicators of similarities or dissimilarities in protein function. We here describe a method to compare the electrostatic properties within protein families objectively and quantitatively. Three-dimensional structures are built from database sequences by comparative modeling. Molecular potentials are then computed for these with a continuum solvation model by finite difference solution of the Poisson-Boltzmann equation or analytically as a multipole expansion that permits rapid comparison of very large datasets. This approach is applied to 104 members of the Pleckstrin homology (PH) domain family. The deviation of the potentials of the homology models from those of the corresponding experimental structures is comparable to the variation of the potential in an ensemble of structures from nuclear magnetic resonance data or between snapshots from a molecular dynamics simulation. For this dataset, the results for analysis of the full electrostatic potential and the analysis using only monopole and dipole terms are very similar. The electrostatic properties of the PH domains are generally conserved despite the extreme sequence divergence in this family. Notable exceptions from this conservation are seen for PH domains linked to a Db1 homology (DH) domain and in proteins with internal PH domain repeats.

Data Interpretation, Statistical↗

Chromosome aberrations in hospital workers: evidence from surveillance studies in Italy (1963-1993).

Hospital workers are occupationally exposed to various agents known or suspected to induce chromosome damage, the most studied being ionizing radiation. To determine the extent of chromosome damage in peripheral blood lymphocytes in this population, taking into account temporal changes and job titles, a re-analysis of cytogenetic studies performed in four Italian laboratories in the period 1965-1993 was carried out. A total of 871 hospital workers and 617 controls, mainly coming from ad hoc studies or surveillance programs in occupational groups potentially exposed to ionizing radiation, were examined. The exposed to controls frequency ratio of chromosome aberrations was evaluated as the measure of effect within each dataset by job title, using multivariate Poisson regression analysis, which allowed an efficient control of confounding. Increased frequency of chromosome-type aberrations among exposed subjects was found in all datasets, especially in those dealing with older data. Significantly higher frequencies are reported for various job titles, particularly for orthopedists, radiologists, anesthesists, and nurses among paramedical occupations. Decrease in exposure to ionizing radiation in hospital workers was documented through a targeted study in the critical group of radiologists. A similar time-related reduction in the frequency of chromosome-type aberrations also has been reported by the surveillance studies carried out over the most recent decades. These data substantiate the use of chromosome-type aberrations as biomarkers of exposure in this occupational setting in the period evaluated. However, the increases observed also in workers with doubtful exposure to ionizing radiation indicate that other chromosome-damaging agents may be involved and, in turn, suggest the extension of surveillance to a larger number of occupations.

Adult↗

Single base-pair substitutions in pathology and evolution: two sides to the same coin.

Relative single base-pair substitution rates in human genes, derived from a collection of > 2,700 point mutations causing human genetic disease, were related to the results of an evolutionary gene/pseudogene comparison. At the mononucleotide level, notable differences between the two datasets were confined to C-to-T and G-to-A transitions, both being rarer in gene/pseudogene alignments than among disease-associated lesions. Relative nearest neighbour-dependent substitution rates were found to be similar in the two datasets, indicating the long-term stability of these parameters during human genome evolution. Allowing for the 5' and 3' nucleotides flanking mutated sites, the primary likelihood of mutation generation could be demonstrated to be biased toward the avoidance of replacements that: (1) change the chemical characteristics of the encoded amino acid residue substantially, and (2) have a high chance of resulting in genetic disease in humans. A similar bias is also reflected in the evolutionary history of human and rodent proteins: amino acid replacements that currently exhibit a high likelihood of coming to clinical attention have been less likely to be accepted during protein evolution.

Base Composition↗

Liver of the "visible man".

Endoscopic surgery, also called minimally invasive surgery, is presumed drastically to reduce postoperative morbidity and thus to offer both human and economic benefits. For the surgeon, however, this approach leads to a number of gestural challenges that require extensive training to be mastered. In order to replace experimentation on animals and patients, we developed a simulator for endoscopic surgery. To achieve this goal, a first step was to develop a working prototype, a "standard patient," on which the informatic and microengineering tools could be validated. We used the visible man dataset for this purpose. The external shape of the visible man's liver, his biliary passages, and his extrahepatic portal system turned out to be fully within the standard pattern of normal anatomy. Anatomic variations were observed in the intrahepatic right portal vein, the hepatic veins, and the arterial blood supply to the liver. Thus, the visible man dataset reveals itself to be well suited for the simulation of minimally invasive surgical operation such as endoscopic cholecystectomy.

Adult↗

Hypescheme: an operational criteria checklist and minimum data set for molecular genetic studies of attention deficit and hyperactivity disorders.

Investigators engaged in mapping the genetic basis of attention deficit hyperactivity disorder (ADHD) currently use a number of measures for the collection of clinical information. This gives rise to difficulties in comparing datasets and research communications between independent groups. This paper describes the development of Hypescheme, which is an operational criteria checklist for ADHD, oppositional defiant disorder (ODD), and conduct disorder (CD), and is proposed as a minimum dataset for those engaged in molecular genetic studies of ADHD. Hypescheme consists of a computerised data checklist system that includes all the operational criteria required for both DSM-IV and ICD-10 diagnostic criteria and a systematic record of information about comorbid psychiatric, developmental, and neurological disorders. Using this data, an algorithm applies both DSM-IV and ICD-10 criteria to generate operational diagnostics under both these systems. Hypescheme is not designed to replace current assessment protocols but to be a final common checklist that can be completed by experienced researchers using all available data.

Attention Deficit Disorder with Hyperactivity↗

Utilization of Alzheimer's disease community resources by Asian-Americans in California.

Alzheimer's disease is as prevalent among Asian ethnic minority groups as among Caucasians. We explored Asian groups' utilization of available Alzheimer's disease services in California, using a uniquely large sample of Asian-Americans. The Minimum Uniform Dataset includes data from nine California Alzheimer's Disease Diagnostic and Treatment Centers. Of the 9,451 cases included in the Minimum Utilizable Dataset, 4.2% were Asian (primarily Chinese), 0.8% Filipino, 0.3% Pacific Islander, and 75.9% Caucasian. In comparison to their numbers within the nine California countries served, Asian ethnic elders were underrepresented in enrollment by approximately 50%, except at one center where all staff were bilingual. The centers referred a significantly greater proportion of Asian than Caucasian patients for financial help (47.8 vs. 7.4%, P < 0.001), case management (47.8 vs. 22.3%, P < 0.001), and to Alzheimer's disease day care (41.3 vs. 28.4%, P < 0.05). A significantly greater proportion of Asian caregivers received referrals to caregiver resource centers (32.6 vs. 61.3%, P < 0.001) and financial help (29.6 vs. 4.7%, P < 0.001). A smaller proportion of Asian patients received referrals to home health services than Caucasians (4.3 vs. 14.9%, P < 0.05). Filipino patients were also referred more frequently to financial assistance than Caucasians (P < 0.05). Asians and Pacific Islanders under-enroll at centers specializing in AD care. Bilingual staff at centers specializing in dementia care, training for community physicians who treat these patients, and establishment of caregiver support groups within Asian and Pacific Islander communities may enhance the enrollment of these elders. AD care centers in areas supporting Asian and Filipino families may need to concentrate resources on providing financial assistance in case management.

Aged↗

Peptide mass fingerprinting peak intensity prediction: extracting knowledge from spectra.

Matrix-assisted laser desorption/ionization-time of flight mass spectrometry has become a valuable tool in proteomics. With the increasing acquisition rate of mass spectrometers, one of the major issues is the development of accurate, efficient and automatic peptide mass fingerprinting (PMF) identification tools. Current tools are mostly based on counting the number of experimental peptide masses matching with theoretical masses. Almost all of them use additional criteria such as isoelectric point, molecular weight, PTMs, taxonomy or enzymatic cleavage rules to enhance prediction performance. However, these identification tools seldom use peak intensities as parameter as there is currently no model predicting the intensities based on the physicochemical properties of peptides. In this work, we used standard datamining methods such as classification and regression methods to find correlations between peak intensities and the properties of the peptides composing a PMF spectrum. These methods were applied on a dataset comprising a series of PMF experiments involving 157 proteins. We found that the C4.5 method gave the more informative results for the classification task (prediction of the presence or absence of a peptide in a spectra) and M5' for the regression methods (prediction of the normalized intensity of a peptide peak). The C4.5 result correctly classified 88% of the theoretical peaks; whereas the M5' peak intensities had a correlation coefficient of 0.6743 with the experimental peak intensities. These methods enabled us to obtain decision and model trees that can be directly used for prediction and identification of PMF results. The work performed permitted to lay the foundations of a method to analyze factors influencing the peak intensity of PMF spectra. A simple extension of this analysis could lead to improve the accuracy of the results by using a larger dataset. Additional peptide characteristics or even PMF experimental parameters can also be taken into account in the datamining process to analyze their influence on the peak intensity. Furthermore, this datamining approach can certainly be extended to the tandem mass spectrometry domain or other mass spectrometry derived methods.

Acrylamide↗

Limitations of stratifying sib-pair data in common disease linkage studies: an example using chromosome 10p14-10q11 in type 1 diabetes.

IDDM10 on chromosome 10p11-q11 has been identified as a putative diabetes susceptibility locus through affected sib-pair (ASP) linkage analysis in UK nuclear families [Davies et al., 1994: Nature 371:130-136; Reed et al., 1997: Hum Mol Genet 6:1011-1016; Mein et al., 1998: Nat Genet 19:297-300]. We extended analysis of linkage to type 1 diabetes in this region by typing a total of 61 markers in a maximum of 418 UK sib-pairs (UK418; peak MLS = 3.84). We then stratified the dataset based on analyses performed previously by both our group [Mein et al., 1998: Nat Genet 19:297-300] and others [Paterson et al., 1999: Hum Hered 49:197-204; Paterson and Petronis, 1999a: Am J Med Genet 84:15-19; Paterson and Petronis, 2000a: J Med Genet 37:186-191; Paterson and Petronis, b: Eur J Hum Genet 8:145-148] and used a permutation procedure to assess the significance of the results. We conclude that the results obtained had a high probability of occurring by chance alone. These data highlight the limitations of stratifying small datasets (n < 500) by additional criteria and the recurrent problems of multiple testing in genetic analysis.

Age Factors↗

CHIP: Defining a dimension of the vulnerability to attention deficit hyperactivity disorder (ADHD) using sibling and individual data of children in a community-based sample.

We are taking a quantitative trait approach to the molecular genetic study of attention deficit hyperactivity disorder (ADHD) using a truncated case-control association design. An epidemiological sample of children aged 5 to 15 years was evaluated for symptoms of ADHD using a parent rating scale. Individuals scoring high or low on this scale were selected for further investigation with additional questionnaires and DNA analysis. Data in studies like this are typically complicated. In the study reported on here, individuals have from 1 to 4 questionnaires completed on them and the sample is composed of a mixture of singletons and siblings. In this paper, we describe how we used a genetic hierarchical model to fit our data, together with a twin dataset, in order to estimate genetic factor loadings. Correlation matrices were estimated for our data using a maximum likelihood approach to account for missing data. We describe how we used these results to create a composite score, the heritability of which was estimated to be acceptably high using the twin dataset. This score measures a quantitative dimension onto which molecular genetic data will be mapped.

Adolescent↗

Role of geographic information systems in birth defects surveillance and research.

BACKGROUND: With the significant advancement of geographic information systems (GIS), mapping and evaluating the spatial distribution of health events has become easier. We examine the role of GIS in birth defects surveillance and research. METHODS: We briefly describe the geocoding process and potential problems in accuracy of the obtained geocodes, and some of the capabilities and limitations of GIS. We illustrate how GIS has been applied using the Metropolitan Atlanta Congenital Defects Program geocoded dataset. We provide some comments on potential data quality and confidentiality issues with birth defects in relation to GIS. RESULTS: It is desirable to geocode addresses using a multistrategy approach to achieve a high-quality and accurate GIS dataset. Beyond the basic but important function of mapping, sophisticated statistical approaches and software are available to analyze the spatial or spatial-temporal occurrence of birth defects, alone or in association with environmental hazards, and to present this information without compromising the confidentiality of the subjects. CONCLUSIONS: We recommend a broad and systematic use of GIS in birth defects spatial surveillance and research.

Congenital Abnormalities↗

The effect of GeneChip gene definitions on the microarray study of cancers.

The Affymetrix GeneChip is a popular microarray platform for genome-wide expression profiling and has been widely used in functional genomics especially in the classification of cancers. Due to the updating of genome data, much of the genome information with which the chips were designed is out-of-date and it has been reported that many of the genes/transcripts on the chips differ from their original definition when mapping the probes to the new genome information. Dai et al. have reported that the updated definition can cause as much as 30-50% discrepancy in the genes selected as differentially expressed on a heart tissue expression profiling dataset. Understanding the nature of this difference is therefore very important for the utilization of the data. In this work, with a large cancer dataset as an example, we compared two major definitions and investigated their effects on classification, clustering, discovery of differentially expressed genes and gene-set-based analysis. Results show that the two definitions agree well on clustering and classification results but genes and gene sets discovered as differentially expressed or enriched can be very different. Discoveries based on the Affymetrix definition can cover most of those based on the new definition, but tend to have more false positives.

Base Sequence↗

Fisher information matrix of the Dirichlet-multinomial distribution.

In this paper we derive explicit expressions for the elements of the exact Fisher information matrix of the Dirichlet-multinomial distribution. We show that exact calculation is based on the beta-binomial probability function rather than that of the Dirichlet-multinomial and this makes the exact calculation quite easy. The exact results are expected to be useful for the calculation of standard errors of the maximum likelihood estimates of the beta-binomial parameters and those of the Dirichlet-multinomial parameters for data that arise in practice in toxicology and other similar fields. Standard errors of the maximum likelihood estimates of the beta-binomial parameters and those of the Dirichlet-multinomial parameters, based on the exact and the asymptotic Fisher information matrix based on the Dirichlet distribution, are obtained for a set of data from Haseman and Soares (1976), a dataset from Mosimann (1962) and a more recent dataset from Chen, Kodell, Howe and Gaylor (1991). There is substantial difference between the standard errors of the estimates based on the exact Fisher information matrix and those based on the asymptotic Fisher information matrix.

Abnormalities, Drug-Induced↗

Bridging the Python Training Gap for Bioscientists in Brazil: Improvements and Challenges.

The rapid evolution of high-throughput technologies in biosciences generates vast and diverse datasets, demanding that bioscientists develop advanced data manipulation and analysis skills. Python, with its versatility and powerful libraries, has become a crucial tool for managing these datasets. However, a significant lack of programming training for bioscientists persists in many countries. To address this knowledge gap in Brazil, the Brazilian Python Workshop for Biological Data was introduced several years ago, focusing on fundamental programming concepts and data handling techniques using popular Python libraries. Despite positive feedback from earlier editions, persistent challenges necessitated continuous adaptation to meet the evolving needs of bioscientists. This work describes the advancements implemented in the 2021 and 2022 editions of the workshop and discusses suggestions for its ongoing enhancement. Key innovations were introduced in the workshop's structure and coordination, including new committees and a code of conduct. Feedback forms were updated for real-time adjustments, and the event's reach was expanded to increase geographical diversity. New didactic strategies, such as pair-teaching, code clubs, and the integration of ICTs, were implemented to enhance learning outcomes. Programming best practices and scientific reproducibility were emphasized through talks and hands-on activities guided by PEP8 conventions. Furthermore, scientific dissemination was intensified through an increased social media presence and participation in international events. Finally, we present updated recommendations for students, researchers, and educators interested in organizing similar initiatives.

Brazil↗

Fast and precise access to enantiomerization rate constants in dynamic chromatography.

An analytical solution of the unified equation to evaluate elution profiles of interconverting enantiomers in dynamic chromatography is presented. Rate constants k1 and k(-1) and Gibbs activation energies are directly obtained from the chromatographic parameters (retention times tR A and tR A of the interconverting enantiomers, the peak widths at half height wA and wB, and the relative plateau height hp), and the initial amounts A0 and B0 of the enantiomers without any iterative and time consuming computational step. Therefore, this equation is no longer limited to racemic analytes. The analytical solution presented here was validated by comparison with a dataset of 125,000 simulated elution profiles of enantiomerizations. Furthermore, it was found that the recovery rate from a defined dataset is on average 40% higher using the unified equation compared to evaluation methods based on iterative computer simulation. The new equation was applied to determine the enantiomerization rate constant of 1-n-butyl-2-tert-butyldiaziridine by enantioselective gas chromatography. The activation parameters (DeltaH(double dagger) = 112.6 +/- 2.5 kJ/mol and DeltaS(double dagger) = -27 +/- 2 J/(K mol) were obtained from temperature-dependent measurements between 100 degrees C and 140 degrees C in 10K steps.

Journal Article↗

A neurocomputational model for prostate carcinoma detection.

BACKGROUND: Current guidelines for prostate carcinoma screening rely primarily on the digital rectal examination (DRE) and prostate specific antigen (PSA). Well described patient risk factors for prostate carcinoma also include age, ethnicity, family history, and complexed PSA. However, due to the nonlinear relation of each of these variables with prostate carcinoma, it is difficult to predict reliably each patient's risk based on linear univariate analysis. The authors investigated a neural network to model the risk of prostate carcinoma by seven readily available clinical features. METHODS: The database for the current study comprised 3268 men recently evaluated for the early detection of prostate carcinoma. The seven clinical features evaluated included age, race, family history, International Prostate Symptom Score (IPSS), DRE, and total and complexed PSA. Three hundred forty-eight subjects in the dataset included men with determined prostate biopsy outcomes and for whom at least 6 of 7 features were available. The dataset was divided randomly into a training set (60%) and a test set (40%), with n1/n2 cross-validation used to evaluate model accuracy, and was modeled with linear and quadratic discriminant function analysis and a neural computational system. After a model with acceptable goodness of fit was achieved, reverse regression analysis using Wilks's generalized likelihood ratio test was performed to evaluate the statistical significance of each input variable. RESULTS: The receiving operating characteristic (ROC) area for the neural computational system in the test set was 0.825, whereas total PSA and complexed PSA alone had ROC areas of 0.678 and 0.697, respectively. The ROC area of logistic regression in the test set was 0.510, linear discriminant function analysis was 0.674, and quadratic discriminant function analysis was 0.011. All were significantly less than the ROC area of the neural computational model (all Ps < 0.002). Reverse regression based on Wilks's generalized likelihood ratio test demonstrated each input feature to be highly significant to the model (all Ps << 0.000001). CONCLUSIONS: The authors modeled a combination of well described patient risk factors for prostate carcinoma using a neural computational system with acceptable goodness of fit. They demonstrated that each of the seven variates on which the model was based was critically significant to model performance. The authors presented this model for clinical use and suggested that clinicians use it in deciding to perform prostate biopsy.

Age Factors↗