Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Dataset”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Gene functional annotation by statistical analysis of biomedical articles.

BACKGROUND: Functional annotation of genes is an important task in biology since it facilitates the characterization of genes relationships and the understanding of biochemical pathways. The various gene functions can be described by standardized and structured vocabularies, called bio-ontologies. The assignment of bio-ontology terms to genes is carried out by means of applying certain methods to datasets extracted from biomedical articles. These methods originate from data mining and machine learning and include maximum entropy or support vector machines (SVM). PURPOSE: The aim of this paper is to propose an alternative to the existing methods for functionally annotating genes. The methodology involves building of classification models, validation and graphical representations of the results and reduction of the dimensions of the dataset. METHODS: Classification models are constructed by Linear discriminant analysis (LDA). The validation of the models is based on statistical analysis and interpretation of the results involving techniques like hold-out samples, test datasets and metrics like confusion matrix, accuracy, recall, precision and F-measure. Graphical representations, such as boxplots, Andrew's curves and scatterplots of the variables resulting from the classification models are also used for validating and interpreting the results. RESULTS: The proposed methodology was applied to a dataset extracted from biomedical articles for 12 Gene Ontology terms. The validation of the LDA models and the comparison with the SVM show that LDA (mean F-measure 75.4%) outperforms the SVM (mean F-measure 68.7%) for the specific data. CONCLUSION: The application of certain statistical methods can be beneficial for functional gene annotation from biomedical articles. Apart from the good performance the results can be interpreted and give insight of the bio-text data structure.

Abstracting and Indexing↗

Models developed by three techniques did not achieve acceptable prediction of binary trauma outcomes.

BACKGROUND AND OBJECTIVES: To develop prediction models for outcomes following trauma that met prespecified performance criteria. To compare three methods of developing prediction models: logistic regression, classification trees, and artificial neural networks. METHODS: Models were developed using a 1996-2001 dataset from a major trauma center in Victoria, Australia. Developed models were subjected to external validation using the first year of data collection, 2001-2002, from a state-wide trauma registry for Victoria. Different authors developed models for each method. All authors were blinded to the validation dataset when developing models. RESULTS: Prediction models were developed for an intensive care unit stay following trauma (prevalence 23%) using information collected at the scene of the injury. None of the three methods gave a model that satisfied the performance criteria of sensitivity >80%, positive predictive value >50% in the validation dataset. Prediction models were also developed for death (prevalence 2.9%) using hospital-collected information. The performance criteria of sensitivity >95%, specificity >20% in the validation dataset were not satisfied by any model. CONCLUSION: No statistical method of model development was optimal. Prespecified performance criteria provide useful guides to interpreting the performance of developed models.

Adult↗

Agri-environmental schemes: their role in reversing floral decline in the Brue floodplain, Somerset, UK.

This paper explores whether the introduction of an agri-environmental scheme has altered the course of long-term trends in plant species abundance in the Somerset Levels and Moors Environmentally Sensitive Area (ESA), UK. A semi-quantitative approach has been taken which integrates disparate but important historical datasets relating to flora and land management with more contemporary digital information. Species datasets from four time periods throughout the 20th century have been collated within a Geographic Information System and analysed with respect to ancillary data relating to elevation, under-drainage and ESA designation. Qualitative reconstruction of the historical ecology of this internationally important area of lowland wet grassland showed that a steady decline in abundance and extent of key components of the flora had already started by 1900. Analysis of historical under-drainage records dating from 1940s to 1980s showed a clear link between the length of time an area had been under-drained and the subsequent diversity of flora recorded in later surveys. In addition, the relative persistence of the rarer components of the wetland flora between surveys in 1980 and 1997 was related to the spatial pattern of under-drainage on the site since 1940. When overall species diversity was compared before and after ESA designation (1980-1997) there was some evidence of an increase in the number of species present and their spatial extent. The historical dataset provided useful contextual information with respect to species trends and allowed the interpretation of contemporary datasets to be placed within a longer timeframe. This pilot study using 18 species gives some evidence that long-established trends in species decline in the Somerset Levels and Moors ESA are starting to be reversed.

Agriculture↗

The phylogenetic position of Morotopithecus.

The phylogenetic relationship of the Ugandan Miocene hominoid Morotopithecus bishopi to fossil and living hominoids remains to be determined. In a cladistic approach to this question, we used three published Miocene character sets as the basis of a phylogenetic analysis: J. Hum. Evol. 29 (1995) 101; Function, Phylogeny, and Fossils: Miocene Hominoid Evolution and Adaptations, 1997, 389. Because these datasets often describe the same anatomy using different characters and states, three different datasets were created to reflect these alternatives. In addition, new postcranial characters describable in Morotopithecus were added to each of the above datasets and a fourth dataset was created using only postcranial characters. The most parsimonious tree(s) recovered in all analyses consistently placed Morotopithecus as a sister taxon to the extant great apes, with Hylobates sister to this clade. Morotopithecus was also consistently more derived than Proconsul, Afropithecus, and Kenyapithecus (as defined prior to the description of Equatorius), but less derived than Oreopithecus, Sivapithecus (only craniodentally) and Dryopithecus. These results imply that Morotopithecus is more derived than Hylobates. However, gibbons are believed to have branched off by at least 18 Ma while Morotopithecus is dated at >20.6 Ma. Possible explanations include: (1) the dating of the Morotopithecus material is too old; (2) the Hylobates divergence time has been underestimated; (3) the great ape condition, and not that of Hylobates, is primitive for hominoids; (4) the similarities of Morotopithecus and great apes are homoplasies. Given current evidence, the first possibility is unlikely, but it is not possible to choose definitively between the latter three possibilities. This conclusion is supported by the fact that despite the consistencies of the analyses, the addition of Morotopithecus and the use of different characters had a large effect on the placement of other Miocene taxa. This raises questions as to the robustness of the connections between Miocene taxa and extant hominoids since different results can be achieved by changing either a few characters, or by adding a single taxon. Many of the characters used to estimate phylogeny may need to be reassessed before a reliable assessment of the phylogenetic position of Morotopithecus can be achieved.

Animals↗

Kinase inhibitor recognition by use of a multivariable QSAR model.

We have applied a retrosynthetic program to determine the scaffold and R-group chemical space seen within a library of known kinase inhibitors and non-kinase drug-like molecules. Comparison of the differences quickly revealed that kinase inhibitors are distinct in several chemical fragment and physical properties. We then applied these descriptors in a multivariable quantitative structure-activity relationship (QSAR) model with the goal to distinguish kinase inhibitors from non-kinase drug-like molecules. This model is heuristic in that it was trained over a dataset of 258 known kinase inhibitors and 230 non-kinase drug molecules. The final model recognized 98% of the training set as being kinase inhibitors and had a false positive rate of 15%. This trait for false positives was accepted out of a desire to maintain diversity and not miss possible good kinase inhibitors for screening. The model was validated by reserving a portion of the datasets as test sets, which were not included in the QSAR model building stage. This was done repetitively for different percentiles of the total dataset population. It was seen that model recognition and false positive were only slightly damaged well down to a 70% reserve (30% dataset used for QSAR model training while 70% used for reserve test set). Beyond 70%, the QSAR models were inconsistent, signifying that the training sets were inadequately diverse to represent the greater reserve test sets. We applied this model to evaluate the commercial kinase libraries available from Asinex, BioFocus, ChemDiv and LifeChemicals to facilitate purchase decisions for compounds for HTS for lead compounds. We observed that there are significant differences in populations of recognizable kinase inhibitors across the vendors analyzed, with BioFocus showing the greatest population of kinase like molecules.

Computational Biology↗

Weighted least-squares deconvolution method for discovery of group differences between complex biofluid 1H NMR spectra.

Biomarker discovery through analysis of high-throughput NMR data is a challenging, time-consuming process due to the requirement of sophisticated, dataset specific preprocessing techniques and the inherent complexity of the data. Here, we demonstrate the use of weighted, constrained least-squares for fitting a linear mixture of reference standard data to complex urine NMR spectra as an automated way of utilizing current assignment knowledge and the ability to deconvolve confounded spectral regions. Following the least-squares fit, univariate statistics were used to identify metabolites associated with group differences. This method was evaluated through applications on simulated datasets and a murine diabetes dataset. Furthermore, we examined the differential ability of various weighting metrics to correctly identify discriminative markers. Our findings suggest that the weighted least-squares approach is effective for identifying biochemical discriminators of varying physiological states. Additionally, the superiority of specific weighting metrics is demonstrated in particular datasets. An additional strength of this methodology is the ability for individual investigators to couple this analysis with laboratory specific preprocessing techniques.

Algorithms↗

A study of spectral integration and normalization in NMR-based metabonomic analyses.

Metabonomics involves the quantitation of the dynamic multivariate metabolic response of an organism to a pathological event or genetic modification [J.K. Nicholson, J.C. Lindon, E. Holmes, Xenobiotica 29 (1999) 1181-1189]. The analysis of these data involves the use of appropriate multivariate statistical methods; Principal Component Analysis (PCA) has been documented as a valuable pattern recognition technique for 1H NMR spectral data [J.T. Brindle, H. Antti, E. Holmes, G. Tranter, J.K. Nicholson, H.W. Bethell, S. Clarke, P.M. Schofield, E. McKilligin, D.E. Mosedale, D.J. Grainger, Nat. Med. 8 (2002) 1439-1444; B.C. Potts, A.J. Deese, G.J. Stevens, M.D. Reily, D.G. Robertson, J. Theiss, J. Pharm. Biomed. Anal. 26 (2001) 463-476; D.G. Robertson, M.D. Reily, R.E. Sigler, D.F. Wells, D.A. Paterson, T.K. Braden, Toxicol. Sci. 57 (2000) 326-337; L.C. Robosky, D.G. Robertson, J.D. Baker, S. Rane, M.D. Reily, Comb. Chem. High Throughput Screen. 5 (2002) 651-662]. Prior to PCA the raw data is typically processed through four steps; (1) baseline correction, (2) endogenous peak removal, (3) integration over spectral regions to reduce the number of variables, and (4) normalization. The effect of the size of spectral integration regions and normalization has not been well studied. The variability structure and classification accuracy on two distinctly different datasets are assessed via PCA and a leave-one-out cross-validation approach under two normalization approaches and an array of spectral integration regions. The first dataset consists of urine from 15 male Wistar-Hannover rats dosed with ANIT measured at five time points, mimicking drug-induced cholangiolitic hepatitis [D.G. Robertson, M.D. Reily, R.E. Sigler, D.F. Wells, D.A. Paterson, T.K. Braden, Toxicol. Sci. 57 (2000) 326-337; J.P. Shockcor, E. Holmes, Curr. Top. Med. Chem. 2 (2002) 35-51; N.J. Waters, E. Holmes, A. Williams, C.J. Waterfield, R.D. Farrant, J.K. Nicholson, Chem. Res. Toxicol. 14 (2001) 1401-1412]. The second data is serum samples from young male C57BL/6 mice subjected to instillation of pancreatic elastase producing emphysema type symptoms [C. Kuhn, S.Y. Yu, M. Chraplyvy, H.E. Linder, R.M. Senior, Lab. Invest. 34 (1976) 372-380; C. Kuhn, R.M. Senior, Lung 155 (1978) 185-197]. This study indicates that independent of the normalization method the classification accuracy achieved from metabonomic studies is not highly sensitive to the size of the spectral integration region. Additionally, both datasets scaled to mean zero and unity variance (auto-scaled) have higher variability within classification accuracy over spectral integration window widths than data scaled to the total intensity of the spectrum. Of the top 10 latent variables for the ANIT dataset the auto-scale normalization has standard deviations larger than the total-scale in seven cases. In the case of the elastase all standard deviations are larger for the auto-scaling.

Animals↗

An introductory practical guide to secondary data analysis in pediatric urology.

INTRODUCTION: Secondary data analysis (SDA) has become an increasingly important approach in pediatric urology, enabling the study of long-term outcomes, care variation, and disparities in populations with chronic or congenital urologic conditions. With the growing availability of large datasets, a structured approach to designing and conducting SDA studies is increasingly relevant. OBJECTIVES: To provide an introductory, practical guide to SDA in pediatric urology by (1) summarizing commonly used data sources with representative studies, (2) outlining a stepwise approach to designing and executing SDA studies, and (3) highlighting key methodological considerations, limitations, and opportunities for future work. STUDY DESIGN: Narrative review of existing literature and commonly used datasets relevant to pediatric urology, including administrative claims, hospital encounter databases, clinical registries, electronic health record networks, and population-based surveys. RESULTS: Data sources differ in scope, clinical granularity, longitudinal follow-up, and representativeness, and each is suited to specific research questions. We present a practical workflow for SDA, including dataset selection, cohort definition, and analytic planning. Linkage across datasets can provide a more comprehensive view of care patterns and outcomes, although feasibility is influenced by legal, technical, and data-quality constraints. DISCUSSION: SDA enables population-level analyses and the study of rare conditions that are challenging to evaluate through single-center or prospective designs. However, careful cohort definition, feasibility assessment, and awareness of data limitations are essential to ensure validity and interpretability. CONCLUSION: SDA provides a scalable, cost-efficient framework for generating meaningful evidence in pediatric urology. Continued efforts to harmonize data elements, improve linkage infrastructure, and support cross-institution collaboration will enhance the quality and impact of future research. This article provides a practical framework and examples to support the design and execution of SDA studies.

Humans↗

Automatic particle selection: results of a comparative study.

Manual selection of single particles in images acquired using cryo-electron microscopy (cryoEM) will become a significant bottleneck when datasets of a hundred thousand or even a million particles are required for structure determination at near atomic resolution. Algorithm development of fully automated particle selection is thus an important research objective in the cryoEM field. A number of research groups are making promising new advances in this area. Evaluation of algorithms using a standard set of cryoEM images is an essential aspect of this algorithm development. With this goal in mind, a particle selection "bakeoff" was included in the program of the Multidisciplinary Workshop on Automatic Particle Selection for cryoEM. Twelve groups participated by submitting the results of testing their own algorithms on a common dataset. The dataset consisted of 82 defocus pairs of high-magnification micrographs, containing keyhole limpet hemocyanin particles, acquired using cryoEM. The results of the bakeoff are presented in this paper along with a summary of the discussion from the workshop. It was agreed that establishing benchmark particles and using bakeoffs to evaluate algorithms are useful in promoting algorithm development for fully automated particle selection, and that the infrastructure set up to support the bakeoff should be maintained and extended to include larger and more varied datasets, and more criteria for future evaluations.

Algorithms↗

A method of focused classification, based on the bootstrap 3D variance analysis, and its application to EF-G-dependent translocation.

The bootstrap-based method for calculation of the 3D variance in cryo-EM maps reconstructed from sets of their projections was applied to a dataset of functional ribosomal complexes containing the Escherichia coli 70S ribosome, tRNAs, and elongation factor G (EF-G). The variance map revealed regions of high variability in the intersubunit space of the ribosome: in the locations of tRNAs, in the putative location of EF-G, and in the vicinity of the L1 protein. This result indicated heterogeneity of the dataset. A method of focused classification was put forward in order to sort out the projection data into approximately homogenous subsets. The method is based on the identification and localization of a region of high variance that a subsequent classification step can be focused on by the use of a 3D spherical mask. After initial classification, template volumes are created and are subsequently refined using a multireference 3D projection alignment procedure. In the application to the ribosome dataset, the two resulting structures were interpreted as resulting from ribosomal complexes with bound EF-G and an empty A site, or, alternatively, from complexes that had no EF-G bound but had both A and P sites occupied by tRNA. The proposed method of focused classification proved to be a successful tool in the analysis of the heterogeneous cryo-EM dataset. The associated calculation of the correlations within the density map confirmed the conformational variability of the complex, which could be interpreted in terms of the ribosomal elongation cycle.

Cryoelectron Microscopy↗

Selecting candidate Neisseria gonorrhoeae strains for oropharyngeal gonorrhoea human challenge: a genomics-based analysis of clinical isolates.

BACKGROUND: Neisseria gonorrhoeae is a human pathogen of major public health importance due to its increasing global prevalence and antimicrobial resistance (AMR). Evidence suggests that oropharyngeal infection plays a key role in N gonorrhoeae transmission and AMR; however, our understanding of oropharyngeal gonorrhoea pathogenesis is poor. A controlled human infection model (CHIM) for oropharyngeal gonorrhoea will improve understanding of infection and accelerate urgently needed novel gonorrhoea prevention and therapeutic strategies. As the first step in the development of this CHIM, we describe a systematic approach to CHIM strain selection that leverages genomics and clinical data. METHODS: In this genomics-based analysis, we applied a systematic N gonorrhoeae challenge strain selection strategy incorporating genomic and clinical data to a primary dataset of clinical isolates of N gonorrhoeae collected from adult patients in Victoria, Australia, between Jan 1 and Dec 31, 2017, and July 1, 2019, and June 30, 2021. This selection strategy used clinical, phenotypic, and genomic characteristics to define a set of eight criteria that aimed to ensure the contemporary global clinical relevance of the candidate strains; select strains that would be applicable for the assessment of current and future gonorrhoea vaccines; and maximise participant safety by reducing the risk of disseminated gonococcal infection and clinically significant AMR. We applied these criteria to our primary dataset to generate a panel of potential challenge strains. From this final dataset of potential challenge strains, we predetermined that we would select up to ten isolates to proceed to the next stage of detailed phenotypic characterisation for final N gonorrhoeae CHIM strain selection. FINDINGS: 5881 isolates comprised the primary dataset. After application of the selection criteria, most of the isolates (5795 [98·6%] of 5881) were excluded, mostly due to having clinically significant AMR and poor contemporary global clinical relevance. The remaining 86 N gonorrhoeae challenge strain candidates comprised five multilocus sequence types and six N gonorrhoeae multiantigen sequence types, many of which were represented by a single isolate. Of these 86 strains, five isolates were selected to maximise coverage of the phylogenetically distinct groups within the 86 candidate challenge strains and ensure representation of strains collected from various anatomical sites. INTERPRETATION: We transparently describe a novel, systematic, and rational genomics-based strategy for oropharyngeal gonorrhoea CHIM strain selection that improves the efficiency and transparency of CHIM strain selection and enables identification of contemporary and clinically relevant potential challenge strains. A final N gonorrhoeae challenge strain will be selected from the subset of five shortlisted candidates after detailed phenotypic assessment. FUNDING: Medical Research Future Fund, Australian National Health and Medical Research Council and Australian Government Research Training Program.

Humans↗

Impact of a changed inundation regime caused by climate change and floodplain rehabilitation on population viability of earthworms in a lower River Rhine floodplain.

River floodplains are dynamic and fertile ecosystems where soil invertebrates such as earthworms can reach high population densities. Earthworms are an important food source for a wide range of organisms including species under conservation such as badgers. Flooding, however, reduces earthworm numbers. Populations recover from cocoons that survive floods. If the period between two floods is too short such that cocoons cannot develop into reproductive adults, populations cannot sustain themselves. Both climate change and floodplain rehabilitation change the flooding frequency affecting earthworm populations. The present paper estimates the influence of climate change and floodplain rehabilitation on the viability of earthworm populations in a Dutch floodplain; the Afferdensche and Deestsche Waarden along the River Waal. This floodplain will be part of major river rehabilitation plans of the Dutch government. In those plans, the floodplain will experience the construction of a secondary channel and the removal of part of its minor embankment. To estimate the impact of these plans and climate change, we used a dataset of daily discharges for 1900-2003 for the River Rhine at the Dutch-German border. We perturbed this dataset to obtain two new datasets under climate change scenarios for 2050 and 2100. From the original and two projected datasets we derived the frequency distributions for the annual periods without inundations for the studied floodplain. We subsequently compared the duration of these inundation-free (dry) periods with the maturation age distribution for L. rubellus as derived from a Dynamic Energy Budget model. This comparison yielded in which parts of our study area and under which climate conditions the populations would still be viable, be able to adapt or become extinct. The results show that climate change has almost no adverse effect on earthworm viability. This is because climate change reduces the flooding frequency during the earthworms growing season. Floodplain rehabilitation, on the other hand, reduces the part of the floodplain area where populations can sustain themselves. Before rehabilitation, only 12% of the floodplain area cannot sustain a viable earthworm population. After rehabilitation, this increases to 59%, 28% of which is due to more frequent flooding. Enhanced exposure to soil contaminants may further suppress earthworm viability. This could frustrate further nature development and the viability of earthworm-dependent species such as the badger (Meles meles) or little owl (Athene noctua vidalli species), which is an objective of the river rehabilitation plans in the Netherlands.

Animals↗

Dissecting the ancient rapid radiation of microgastrine wasp genera using additional nuclear genes.

Previous estimates of a generic level phylogeny for the ubiquitous parasitoid wasp subfamily Microgastrinae (Hymenoptera) have been problematic due to short internal branches deep in the phylogeny. These short branches might be attributed to a rapid radiation among the taxa, the use of genes that are unsuitable for the levels of divergence being examined, or insufficient quantity of data. We added over 1200 nucleotides from four nuclear genes to a dataset derived from three genes to produce a dataset of over 3000 nucleotides per taxon. While the number of well-supported short branches in the phylogeny increased, we still did not obtain strong bootstrap support for every node. Parametric and nonparametric bootstrap simulations projected that an enormous, and likely unobtainable, amount of data would be required to get bootstrap support greater than 50% for every node. However, a marked increase in the number of well-supported nodes was seen when we conducted a Bayesian analysis of a combined dataset generated from morphological characters added to the seven gene dataset. Our results suggest that, in some cases, combining morphological and genetic characters may be the most practical way to increase support for short branches deep in a phylogeny.

Animals↗

The genetic association between Cathepsin D and Alzheimer's disease.

The aspartyl protease Cathepsin D has previously been suggested to play a role in the Alzheimer's disease (AD) process because of its ability to cleave the beta-amyloid precursor protein and the possibility that it may be one of the 'secretase' enzymes. A functional C-->T polymorphism in the Cathepsin D gene (CATD) has been reported to be associated with increased risk for AD in Caucasian case-control studies; specifically, the T-carrying genotypes confer increased risk. We have examined this association in our own Caucasian dataset of 210 AD cases and 120 controls, and in an additional Hispanic dataset comprising 79 AD cases and 112 controls. In Hispanics we find a modest interaction between CATD genotype and age of onset on risk for AD, such that the non-T-carrying genotype confers increased risk. In our Caucasian dataset we find no evidence for association between the CATD polymorphism and AD, although we do observe a small tendency towards an increase in the T-carrying genotypes in the case group, consistent with previous studies. We conducted an aggregate analysis of the published Caucasian datasets and found evidence that this CATD polymorphism (or another locus in linkage disequilibrium) does contribute significant, but small (<2%) risk for AD.

Aged↗

Development and validation of the Ontario acute myocardial infarction mortality prediction rules.

OBJECTIVES: To develop and validate simple statistical models that can be used with hospital discharge administrative databases to predict 30-day and one-year mortality after an acute myocardial infarction (AMI). BACKGROUND: There is increasing interest in developing AMI "report cards" using population-based hospital discharge databases. However, there is a lack of simple statistical models that can be used to adjust for regional and interinstitutional differences in patient case-mix. METHODS: We used linked administrative databases on 52,616 patients having an AMI in Ontario, Canada, between 1994 and 1997 to develop logistic regression statistical models to predict 30-day and one-year mortality after an AMI. These models were subsequently validated in two external cohorts of AMI patients derived from administrative datasets from Manitoba, Canada, and California, U.S. RESULTS: The 11-variable Ontario AMI mortality prediction rules accurately predicted mortality with an area under the receiver operating characteristic (ROC) curve of 0.78 for 30-day mortality and 0.79 for one-year mortality in the Ontario dataset from which they were derived. In an independent validation dataset of 4,836 AMI patients from Manitoba, the ROC areas were 0.77 and 0.78, respectively. In a second validation dataset of 112,234 AMI patients from California, the ROC areas were 0.77 and 0.78 respectively. CONCLUSIONS: The Ontario AMI mortality prediction rules predict quite accurately 30-day and one-year mortality after an AMI in linked hospital discharge databases of AMI patients from Ontario, Manitoba and California. These models may also be useful to outcomes and quality measurement researchers in other jurisdictions.

Aged↗

Dynamical components analysis of fMRI data through kernel PCA.

In parallel with standard model-based methods for the analysis of fMRI data, exploratory methods--such as PCA, ICA, and clustering--have been developed to give an account of the dataset with minimal priors: no assumption is made on the data content itself, but the data structure is assumed to show some properties (decorrelation, independence) that allow for the detection of structures of interest. In this paper, we present an alternative that tries to take into account some relevant knowledge for the analysis of the dataset, e.g., the experimental paradigm, while keeping the flexibility of exploratory methods: we use a prior temporal modeling of the data that characterizes each voxel time course. Two implementations are proposed: one based on the General Linear Model, the other one on more flexible short-term predictors, whose complexity is controlled by a Minimum Description Length approach. However, our main concern here is the construction of a multivariate model; the latter is performed with the help of a kernel PCA method that builds a redundant representation of the data through the nonlinearity of the kernel. This allows for a refinement in the description of the (temporal) patterns of interest. In particular, this helps in the characterization of subtle variations in the response to different experimental conditions. We illustrate the usefulness of nonlinearity through the analysis of a synthetic dataset and show on a real dataset how it helps to interpret the experimental results.

Algorithms↗

Mitochondrial DNA, morphology, and the phylogenetic relationships of Antarctic icefishes (Notothenioidei: Channichthyidae).

The Channichthyidae is a lineage of 16 species in the Notothenioidei, a clade of fishes that dominate Antarctic near-shore marine ecosystems with respect to both diversity and biomass. Among four published studies investigating channichthyid phylogeny, no two have produced the same tree topology, and no published study has investigated the degree of phylogenetic incongruence between existing molecular and morphological datasets. In this investigation we present an analysis of channichthyid phylogeny using complete gene sequences from two mitochondrial genes (ND2 and 16S) sampled from all recognized species in the clade. In addition, we have scored all 58 unique morphological characters used in three previous analyses of channichthyid phylogenetic relationships. Data partitions were analyzed separately to assess the amount of phylogenetic resolution provided by each dataset, and phylogenetic incongruence among data partitions was investigated using incongruence length difference (ILD) tests. We utilized a parsimony-based version of the Shimodaira-Hasegawa test to determine if alternative tree topologies are significantly different from trees resulting from maximum parsimony analysis of the combined partition dataset. Our results demonstrate that the greatest phylogenetic resolution is achieved when all molecular and morphological data partitions are combined into a single maximum parsimony analysis. Also, marginal to insignificant incongruence was detected among data partitions using the ILD. Maximum parsimony analysis of all data partitions combined results in a single tree, and is a unique hypothesis of phylogenetic relationships in the Channichthyidae. In particular, this hypothesis resolves the phylogenetic relationships of at least two species (Channichthys rhinoceratus and Chaenocephalus aceratus), for which there was no consensus among the previous phylogenetic hypotheses. The combined data partition dataset provides substantial statistical power to discriminate among alternative hypotheses of channichthyid relationships. These findings suggest the optimal strategy for investigating the phylogenetic relationships of channichthyids is one that uses all available phylogenetic data in analyses of combined data partitions.

Animals↗

Intrinsic disorder in transcription factors.

Intrinsic disorder (ID) is highly abundant in eukaryotes, which reflect the greater need for disorder-associated signaling and transcriptional regulation in nucleated cells. Although several well-characterized examples of intrinsically disordered proteins in transcriptional regulation have been reported, no systematic analysis has been reported so far. To test for the general prevalence of intrinsic disorder in transcriptional regulation, we used the predictor of natural disorder regions (PONDR) to analyze the abundance of intrinsic disorder in three transcription factor datasets and two control sets. This analysis revealed that from 94.13 to 82.63% of transcription factors possess extended regions of intrinsic disorder, relative to 54.51 and 18.64% of the proteins in two control datasets, which indicates the significant prevalence of intrinsic disorder in transcription factors. This propensity of transcription factors to intrinsic disorder was confirmed by cumulative distribution function analysis and charge-hydropathy plots. The amino acid composition analysis showed that all three transcription factor datasets were substantially depleted in order-promoting residues and significantly enriched in disorder-promoting residues. Our analysis of the distribution of disorder within the transcription factor datasets revealed that (a) the AT-hooks and basic regions of transcription factor DNA-binding domains are highly disordered; (b) the degree of disorder in transcription factor activation regions is much higher than that in DNA-binding domains; (c) the degree of disorder is significantly higher in eukaryotic transcription factors than in prokaryotic transcription factors; and (d) the level of alpha-MoRF (molecular recognition feature) prediction is much higher in transcription factors. Overall, our data reflected the fact that eukaryotes with well-developed gene transcription machinery require transcription factor flexibility to be more efficient.

Amino Acid Sequence↗