Search PubMed⌕ Search

Biomedical subjects

Xiaohui Xie

Publications and source records attributed to Xiaohui Xie.

18 recordsLinked to original sources

A family of conserved noncoding elements derived from an ancient transposable element.

The evolutionary origin of the conserved noncoding elements (CNEs) in the human genome remains poorly understood but may hold important clues to their biological functions. Here, we report the discovery of a CNE family with approximately 124 instances in the human genome that demonstrates a clear signature of having been derived from an ancient transposon. The CNE family is also present in the chicken genome, although typically not at orthologous locations. The CNE family is closely related to the active transposon SINE3 in zebrafish and also to a previously uncharacterized transposon in the coelacanth, the so-called "living fossil" belonging to the lobe-finned fish lineage. The mammal, bird, zebrafish, and coelacanth families all share a highly similar core element of approximately 180 bp but have important differences in their 5' and 3' ends. The core element has thus been preserved over 450 million years of evolution, implying an important biological function. In addition, we identify 95 additional CNE families that likely predate the mammalian radiation. The results highlight both the creative role of transposons and the importance of CNE families.

Animals↗

A molecular-properties-based approach to understanding PDZ domain proteins and PDZ ligands.

PDZ domain-containing proteins and their interaction partners are mutated in numerous human diseases and function in complexes regulating epithelial polarity, ion channels, cochlear hair cell development, vesicular sorting, and neuronal synaptic communication. Among several properties of a collection of documented PDZ domain-ligand interactions, we discovered embedded in a large-scale expression data set the existence of a significant level of co-regulation between PDZ domain-encoding genes and these ligands. From this observation, we show how integration of expression data, a comparative genomics catalog of 899 mammalian genes with conserved PDZ-binding motifs, phylogenetic analysis, and literature mining can be utilized to infer PDZ complexes. Using molecular studies we map novel interaction partners for the PDZ proteins DLG1 and CARD11. These results provide insight into the diverse roles of PDZ-ligand complexes in cellular signaling and provide a computational framework for the genome-wide evaluation of PDZ complexes.

Amino Acid Motifs↗

A bivalent chromatin structure marks key developmental genes in embryonic stem cells.

The most highly conserved noncoding elements (HCNEs) in mammalian genomes cluster within regions enriched for genes encoding developmentally important transcription factors (TFs). This suggests that HCNE-rich regions may contain key regulatory controls involved in development. We explored this by examining histone methylation in mouse embryonic stem (ES) cells across 56 large HCNE-rich loci. We identified a specific modification pattern, termed "bivalent domains," consisting of large regions of H3 lysine 27 methylation harboring smaller regions of H3 lysine 4 methylation. Bivalent domains tend to coincide with TF genes expressed at low levels. We propose that bivalent domains silence developmental genes in ES cells while keeping them poised for activation. We also found striking correspondences between genome sequence and histone methylation in ES cells, which become notably weaker in differentiated cells. These results highlight the importance of DNA sequence in defining the initial epigenetic landscape and suggest a novel chromatin-based mechanism for maintaining pluripotency.

Animals↗

A mammalian organelle map by protein correlation profiling.

Protein localization to membrane-enclosed organelles is a central feature of cellular organization. Using protein correlation profiling, we have mapped 1,404 proteins to ten subcellular locations in mouse liver, and these correspond with enzymatic assays, marker protein profiles, and confocal microscopy. These localizations allowed assessment of the specificity in published organellar proteomic inventories and demonstrate multiple locations for 39% of all organellar proteins. Integration of proteomic and genomic data enabled us to identify networks of coexpressed genes, cis-regulatory motifs, and putative transcriptional regulators involved in organelle biogenesis. Our analysis ties biochemistry, cell biology, and genomics into a common framework for organelle analysis.

Animals↗

Systematic identification of human mitochondrial disease genes through integrative genomics.

The majority of inherited mitochondrial disorders are due to mutations not in the mitochondrial genome (mtDNA) but rather in the nuclear genes encoding proteins targeted to this organelle. Elucidation of the molecular basis for these disorders is limited because only half of the estimated 1,500 mitochondrial proteins have been identified. To systematically expand this catalog, we experimentally and computationally generated eight genome-scale data sets, each designed to provide clues as to mitochondrial localization: targeting sequence prediction, protein domain enrichment, presence of cis-regulatory motifs, yeast homology, ancestry, tandem-mass spectrometry, coexpression and transcriptional induction during mitochondrial biogenesis. Through an integrated analysis we expand the collection to 1,080 genes, which includes 368 novel predictions with a 10% estimated false prediction rate. By combining this expanded inventory with genetic intervals linked to disease, we have identified candidate genes for eight mitochondrial disorders, leading to the discovery of mutations in MPV17 that result in hepatic mtDNA depletion syndrome. The integrative approach promises to better define the role of mitochondria in both rare and common human diseases.

Base Sequence↗

A large family of ancient repeat elements in the human genome is under strong selection.

Although conserved noncoding elements (CNEs) constitute the majority of sequences under purifying selection in the human genome, they remain poorly understood. CNEs seem to be largely unique, with no large families of similar elements reported to date. Here, we search for CNEs among the ancestral repeat classes in the human genome and report the discovery of a large CNE family containing >900 members. This family belongs to the MER121 class of repeats. Although the MER121 family members show considerable sequence variation among one another, the individual copies show striking conservation in orthologous locations across the human, dog, mouse, and rat genomes. The element is also present and conserved in orthologous locations in the marsupial, but its genome-wide dispersal postdates the divergence from birds. The comparative genomic data indicate that MER121 does not encode a family of either protein-coding or RNA genes. Although the precise function of these elements remains unknown, the evidence suggests that this unusual family may play a cis-regulatory or structural role in mammalian genomes.

Animals↗

Comparative sequence analysis reveals an intricate network among REST, CREB and miRNA in mediating neuronal gene expression.

BACKGROUND: Two distinct classes of regulators have been implicated in regulating neuronal gene expression and mediating neuronal identity: transcription factors such as REST/NRSF (RE1 silencing transcription factor) and CREB (cAMP response element-binding protein), and microRNAs (miRNAs). How these two classes of regulators act together to mediate neuronal gene expression is unclear. RESULTS: Using comparative sequence analysis, here we report the identification of 895 sites (NRSE) as the putative targets of REST. A set of the identified NRSE sites is present in the vicinity of the miRNA genes that are specifically expressed in brain-related tissues, suggesting the transcriptional regulation of these miRNAs by REST. We have further identified target genes of these miRNAs, and discovered that REST and its cofactor complex are targets of multiple brain-related miRNAs including miR-124a, miR-9 and miR-132. Given the role of both REST and miRNA as repressors, these findings point to a double-negative feedback loop between REST and the miRNAs in stabilizing and maintaining neuronal gene expression. Additionally, we find that the brain-related miRNA genes are highly enriched with evolutionarily conserved cAMP response elements (CRE) in their regulatory regions, implicating the role of CREB in the positive regulation of these miRNAs. CONCLUSION: The expression of neuronal genes and neuronal identity are controlled by multiple factors, including transcriptional regulation through REST and post-transcriptional modification by several brain-related miRNAs. We demonstrate that these different levels of regulation are coordinated through extensive feedbacks, and propose a network among REST, CREB proteins and the brain-related miRNAs as a robust program for mediating neuronal gene expression.

Cyclic AMP Response Element-Binding Protein↗

Genome sequence, comparative analysis and haplotype structure of the domestic dog.

Here we report a high-quality draft genome sequence of the domestic dog (Canis familiaris), together with a dense map of single nucleotide polymorphisms (SNPs) across breeds. The dog is of particular interest because it provides important evolutionary information and because existing breeds show great phenotypic diversity for morphological, physiological and behavioural traits. We use sequence comparison with the primate and rodent lineages to shed light on the structure and evolution of genomes and genes. Notably, the majority of the most highly conserved non-coding sequences in mammalian genomes are clustered near a small subset of genes with important roles in development. Analysis of SNPs reveals long-range haplotypes across the entire dog genome, and defines the nature of genetic diversity within and across breeds. The current SNP map now makes it possible for genome-wide association studies to identify genes responsible for diseases and traits, with important consequences for human and companion animal health.

Animals↗

Systematic discovery of regulatory motifs in human promoters and 3' UTRs by comparison of several mammals.

Comprehensive identification of all functional elements encoded in the human genome is a fundamental need in biomedical research. Here, we present a comparative analysis of the human, mouse, rat and dog genomes to create a systematic catalogue of common regulatory motifs in promoters and 3' untranslated regions (3' UTRs). The promoter analysis yields 174 candidate motifs, including most previously known transcription-factor binding sites and 105 new motifs. The 3'-UTR analysis yields 106 motifs likely to be involved in post-transcriptional regulation. Nearly one-half are associated with microRNAs (miRNAs), leading to the discovery of many new miRNA genes and their likely target genes. Our results suggest that previous estimates of the number of human miRNA genes were low, and that miRNAs regulate at least 20% of human genes. The overall results provide a systematic view of gene regulation in the human, which will be refined as additional mammalian genomes become available.

3' Untranslated Regions↗

Disease gene discovery through integrative genomics.

The availability of complete genome sequences and the wealth of large-scale biological data sets now provide an unprecedented opportunity to elucidate the genetic basis of rare and common human diseases. Here we review some of the emerging genomics technologies and data resources that can be used to infer gene function to prioritize candidate genes. We then describe some computational strategies for integrating these large-scale data sets to provide more faithful descriptions of gene function, and how such approaches have recently been applied to discover genes underlying Mendelian disorders. Finally, we discuss future prospects and challenges for using integrative genomics to systematically discover not only single genes but also entire gene networks that underlie and modify human disease.

Databases, Genetic↗

Learning curves for stochastic gradient descent in linear feedforward networks.

Gradient-following learning methods can encounter problems of implementation in many applications, and stochastic variants are sometimes used to overcome these difficulties. We analyze three online training methods used with a linear perceptron: direct gradient descent, node perturbation, and weight perturbation. Learning speed is defined as the rate of exponential decay in the learning curves. When the scalar parameter that controls the size of weight updates is chosen to maximize learning speed, node perturbation is slower than direct gradient descent by a factor equal to the number of output units; weight perturbation is slower still by an additional factor equal to the number of input units. Parallel perturbation allows faster learning than sequential perturbation, by a factor that does not depend on network size. We also characterize how uncertainty in quantities used in the stochastic updates affects the learning curves. This study suggests that in practice, weight perturbation may be slow for large networks, and node perturbation can have performance comparable to that of direct gradient descent when there are few output units. However, these statements depend on the specifics of the learning problem, such as the input distribution and the target function, and are not universally applicable.

Neural Networks, Computer↗

Learning in neural networks by reinforcement of irregular spiking.

Artificial neural networks are often trained by using the back propagation algorithm to compute the gradient of an objective function with respect to the synaptic strengths. For a biological neural network, such a gradient computation would be difficult to implement, because of the complex dynamics of intrinsic and synaptic conductances in neurons. Here we show that irregular spiking similar to that observed in biological neurons could be used as the basis for a learning rule that calculates a stochastic approximation to the gradient. The learning rule is derived based on a special class of model networks in which neurons fire spike trains with Poisson statistics. The learning is compatible with forms of synaptic dynamics such as short-term facilitation and depression. By correlating the fluctuations in irregular spiking with a reward signal, the learning rule performs stochastic gradient ascent on the expected reward. It is applied to two examples, learning the XOR computation and learning direction selectivity using depressing synapses. We also show in simulation that the learning rule is applicable to a network of noisy integrate-and-fire neurons.

Action Potentials↗

Erralpha and Gabpa/b specify PGC-1alpha-dependent oxidative phosphorylation gene expression that is altered in diabetic muscle.

Recent studies have shown that genes involved in oxidative phosphorylation (OXPHOS) exhibit reduced expression in skeletal muscle of diabetic and prediabetic humans. Moreover, these changes may be mediated by the transcriptional coactivator peroxisome proliferator-activated receptor gamma coactivator-1alpha (PGC-1alpha). By combining PGC-1alpha-induced genome-wide transcriptional profiles with a computational strategy to detect cis-regulatory motifs, we identified estrogen-related receptor alpha (Erralpha) and GA repeat-binding protein alpha as key transcription factors regulating the OXPHOS pathway. Interestingly, the genes encoding these two transcription factors are themselves PGC-1alpha-inducible and contain variants of both motifs near their promoters. Cellular assays confirmed that Erralpha and GA-binding protein a partner with PGC-1alpha in muscle to form a double-positive-feedback loop that drives the expression of many OXPHOS genes. By using a synthetic inhibitor of Erralpha, we demonstrated its key role in PGC-1alpha-mediated effects on gene regulation and cellular respiration. These results illustrate the dissection of gene regulatory networks in a complex mammalian system, elucidate the mechanism of PGC-1alpha action in the OXPHOS pathway, and suggest that Erralpha agonists may ameliorate insulin-resistance in individuals with type 2 diabetes mellitus.

Animals↗

Equivalence of backpropagation and contrastive Hebbian learning in a layered network.

Backpropagation and contrastive Hebbian learning are two methods of training networks with hidden neurons. Backpropagation computes an error signal for the output neurons and spreads it over the hidden neurons. Contrastive Hebbian learning involves clamping the output neurons at desired values and letting the effect spread through feedback connections over the entire network. To investigate the relationship between these two forms of learning, we consider a special case in which they are identical: a multilayer perceptron with linear output units, to which weak feedback connections have been added. In this case, the change in network state caused by clamping the output neurons turns out to be the same as the error signal spread by backpropagation, except for a scalar prefactor. This suggests that the functionality of backpropagation can be realized alternatively by a Hebbian-type learning algorithm, which is suitable for implementation in biological networks.

Algorithms↗

Double-ring network model of the head-direction system.

In the head-direction system, the orientation of an animal's head in space is encoded internally by persistent activities of a pool of cells whose firing rates are tuned to the animal's directional heading. To maintain an accurate representation of the heading information when the animal moves, the system integrates horizontal angular head-velocity signals from the vestibular nuclei and updates the representation of directional heading. The integration is a difficult process, given that head velocities can vary over a large range and the neural system is highly nonlinear. Previous models of integration have relied on biologically unrealistic mechanisms, such as instantaneous changes in synaptic strength, or very fast synaptic dynamics. In this paper, we propose a different integration model with two populations of neurons, which performs integration based on the differential input of the vestibular nuclei to these two populations. We mathematically analyze the dynamics of the model and demonstrate that with carefully tuned synaptic connections it can accurately integrate a large range of the vestibular input, with potentially slow synapses.

Animals↗

Nonlinear dynamics of direction-selective recurrent neural media.

The direction selectivity of cortical neurons can be accounted for by asymmetric lateral connections. Such lateral connectivity leads to a network dynamics with characteristic properties that can be exploited for distinguishing in neurophysiological experiments this mechanism for direction selectivity from other possible mechanisms. We present a mathematical analysis for a class of direction-selective neural models with asymmetric lateral connections. Contrasting with earlier theoretical studies that have analyzed approximations of the network dynamics by neglecting nonlinearities using methods from linear systems theory, we study the network dynamics with nonlinearity taken into consideration. We show that asymmetrically coupled networks can stabilize stimulus-locked traveling pulse solutions that are appropriate for the modeling of the responses of direction-selective neurons. In addition, our analysis shows that outside a certain regime of stimulus speeds the stability of these solutions breaks down, giving rise to lurching activity waves with specific spatiotemporal periodicity. These solutions, and the bifurcation by which they arise, cannot be easily accounted for by classical models for direction selectivity.

Animals↗

Selectively grouping neurons in recurrent networks of lateral inhibition.

Winner-take-all networks have been proposed to underlie many of the brain's fundamental computational abilities. However, not much is known about how to extend the grouping of potential winners in these networks beyond single neuron or uniformly arranged groups of neurons. We show that competition between arbitrary groups of neurons can be realized by organizing lateral inhibition in linear threshold networks. Given a collection of potentially overlapping groups (with the exception of some degenerate cases), the lateral inhibition results in network dynamics such that any permitted set of neurons that can be coactivated by some input at a stable steady state is contained in one of the groups. The information about the input is preserved in this operation. The activity level of a neuron in a permitted set corresponds to its stimulus strength, amplified by some constant. Sets of neurons that are not part of a group cannot be coactivated by any input at a stable steady state. We analyze the storage capacity of such a network for random groups--the number of random groups the network can store as permitted sets without creating too many spurious ones. In this framework, we calculate the optimal sparsity of the groups (maximizing group entropy). We find that for dense inputs, the optimal sparsity is unphysiologically small. However, when the inputs and the groups are equally sparse, we derive a more plausible optimal sparsity. We believe our results are the first steps toward attractor theories in hybrid analog-digital networks.

Brain↗

Threshold behaviour of the maximum likelihood method in population decoding.

We study the performance of the maximum likelihood (ML) method in population decoding as a function of the population size. Assuming uncorrelated noise in neural responses, the ML performance, quantified by the expected square difference between the estimated and the actual quantity, follows closely the optimal Cramer-Rao bound, provided that the population size is sufficiently large. However, when the population size decreases below a certain threshold, the performance of the ML method undergoes a rapid deterioration, experiencing a large deviation from the optimal bound. We explain the cause of such threshold behaviour, and present a phenomenological approach for estimating the threshold population size, which is found to be linearly proportional to the inverse of the square of the system's signal-to-noise ratio. If the ML method is used by neural systems, we expect the number of neurons involved in population coding to be above this threshold.

Artifacts↗