Search PubMed⌕ Search

Biomedical subjects

Ning Lan

Publications and source records attributed to Ning Lan.

11 recordsLinked to original sources

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗

Strategies for structural proteomics of prokaryotes: Quantifying the advantages of studying orthologous proteins and of using both NMR and X-ray crystallography approaches.

Only about half of non-membrane-bound proteins encoded by either bacterial or archaeal genomes are soluble when expressed in Escherichia coli (Yee et al., Proc Natl Acad Sci USA 2002;99:1825-1830; Christendat et al., Prog Biophys Mol Biol 200;73:339-345). This property limits genome-scale functional and structural proteomics studies, which depend on having a recombinant, soluble version of each protein. An emerging strategy to increase the probability of deriving a soluble derivative of a protein is to study different sequence homologues of the same protein, including representatives from thermophilic organisms, based on the assumption that the stability of these proteins will facilitate structural analysis. To estimate the relative merits of this strategy, we compared the recombinant expression, solubility, and suitability for structural analysis by NMR and/or X-ray crystallography for 68 pairs of homologous proteins from E. coli and Thermotoga maritima. A sample suitable for structural studies was obtained for 62 of the 68 pairs of homologs under standardized growth and purification procedures. Fourteen (eight E. coli and six T. maritima proteins) samples generated NMR spectra of a quality suitable for structure determination and 30 (14 E. coli and 16 T. maritima proteins) samples formed crystals. Only three (one E. coli and two T. maritima proteins) samples both crystallized and had excellent NMR properties. The conclusions from this work are: (1) The inclusion of even a single ortholog of a target protein increases the number of samples for structural studies almost twofold; (2) there was no clear advantage to the use of thermophilic proteins to generate samples for structural studies; and (3) for the small proteins analyzed here, the use of both NMR and crystallography approaches almost doubled the number of samples for structural studies.

Archaeal Proteins↗

Ontologies for proteomics: towards a systematic definition of structure and function that scales to the genome level.

A principal aim of post-genomic biology is elucidating the structures, functions and biochemical properties of all gene products in a genome. However, to adequately comprehend such a large amount of information we need new descriptions of proteins that scale to the genomic level. In short, we need a unified ontology for proteomics. Much progress has been made towards this end, including a variety of approaches to systematic structural and functional classification and initial work towards developing standardized, unified descriptions for protein properties. In relation to function, there is a particularly great diversity of approaches, involving placing a protein in structured hierarchies or more-generalized networks and a recent approach based on circumscribing a protein's function through systematic enumeration of molecular interactions.

Computational Biology↗

Efficient and specific repair of sickle beta-globin RNA by trans-splicing ribozymes.

Previously we demonstrated that a group I ribozyme can perform trans-splicing to repair sickle beta-globin transcripts upon transfection of in vitro transcribed ribozyme into mammalian cells. Here, we sought to develop expression cassettes that would yield high levels of active ribozyme after gene transfer. Our initial expression constructs were designed to generate trans-slicing ribozymes identical to those used in our previous RNA transfection studies with ribozymes containing 6-nucleotide long internal guide sequences. The ribozymes expressed from these cassettes, however, were found to be unable to repair sickle beta-globin RNAs. Further experiments revealed that two additional structural elements are important for ribozyme-mediate RNA repair: the P10 interaction formed between the 5' end of the ribozyme and the beginning of the 3' exon and an additional base-pairing interaction formed between an extended guide sequence and the substrate RNA. These optimized expression cassettes yield ribozymes that are able to amend 10%-50% of the sickle beta-globin RNAs in transfected mammalian cells. Finally, a ribozyme with a 5-bp extended guide sequence preferentially reacts with sickle beta-globin RNAs over wild-type beta-globin RNAs, although the wild-type beta-globin transcript forms only a single mismatch with the ribozyme. These results demonstrate that trans-splicing ribozyme expression cassettes can be generated to yield ribozymes that can repair a clinically relevant fraction of sickle beta-globin RNAs in mammalian cells with greatly improved specificity.

Anemia, Sickle Cell↗

Functional profiling of the Saccharomyces cerevisiae genome.

Determining the effect of gene deletion is a fundamental approach to understanding gene function. Conventional genetic screens exhibit biases, and genes contributing to a phenotype are often missed. We systematically constructed a nearly complete collection of gene-deletion mutants (96% of annotated open reading frames, or ORFs) of the yeast Saccharomyces cerevisiae. DNA sequences dubbed 'molecular bar codes' uniquely identify each strain, enabling their growth to be analysed in parallel and the fitness contribution of each gene to be quantitatively assessed by hybridization to high-density oligonucleotide arrays. We show that previously known and new genes are necessary for optimal growth under six well-studied conditions: high salt, sorbitol, galactose, pH 8, minimal medium and nystatin treatment. Less than 7% of genes that exhibit a significant increase in messenger RNA expression are also required for optimal growth in four of the tested conditions. Our results validate the yeast gene-deletion collection as a valuable resource for functional genomics.

Cell Size↗

A small reservoir of disabled ORFs in the yeast genome and its implications for the dynamics of proteome evolution.

We surveyed the sequenced Saccharomyces cerevisiae genome (strain S288C) comprehensively for open reading frames (ORFs) that could encode full-length proteins but contain obvious mid-sequence disablements (frameshifts or premature stop codons). These pseudogenic features are termed disabled ORFs (dORFs). Using homology to annotated yeast ORFs and non-yeast proteins plus a simple region extension procedure, we have found 183 dORFs. Combined with the 38 existing annotations for potential dORFs, we have a total pool of up to 221 dORFs, corresponding to less than approximately 3% of the proteome. Additionally, we found 20 pairs of annotated ORFs for yeast that could be merged into a single ORF (termed a mORF) by read-through of the intervening stop codon, and may comprise a complete ORF in other yeast strains. Focussing on a core pool of 98 dORFs with a verifying protein homology, we find that most dORFs are substantially decayed, with approximately 90% having two or more disablements, and approximately 60% having four or more. dORFs are much more yeast-proteome specific than live yeast genes (having about half the chance that they are related to a non-yeast protein). They show a dramatically increased density at the telomeres of chromosomes, relative to genes. A microarray study shows that some dORFs are expressed even though they carry multiple disablements, and thus may be more resistant to nonsense-mediated decay. Many of the dORFs may be involved in responding to environmental stresses, as the largest functional groups include growth inhibition, flocculation, and the SRP/TIP1 family. Our results have important implications for proteome evolution. The characteristics of the dORF population suggest the sorts of genes that are likely to fall in and out of usage (and vary in copy number) in a strain-specific way and highlight the role of subtelomeric regions in engendering this diversity. Our results also have important implications for the effects of the [PSI+] prion. The dORFs disabled by only a single stop and the mORFs (together totalling 35) provide an estimate for the extent of the sequence population that can be resurrected readily through the demonstrated ability of the [PSI+] prion to cause nonsense-codon read-through. Also, the dORFs and mORFs that we find have properties (e.g. growth inhibition, flocculation, vanadate resistance, stress response) that are potentially related to the ability of [PSI+] to engender substantial phenotypic variation in yeast strains under different environmental conditions. (See genecensus.org/pseudogene for further information.)

Chromosomes, Fungal↗

The importance of biomechanics.

When neuroscientists gather to discuss "Movement and Sensation", they tend to discuss neurons rather than muscles and bones. Neurons may be more interesting, but their roles in motor control depend on the mechanical properties of the system to be controlled. Understanding of those properties has been surprisingly elusive, despite the well-developed disciplines of biomechanics and muscle physiology. Each experimental field has its favorite, often unique preparation. Mathematical models range in scale from individual cross-bridges to articulated limbs, usually written in different computer languages. The shortcomings of such fragmented knowledge become particularly apparent when biomedical engineers must design safe and effective control systems for real limbs, such as for functional electrical stimulation (FES) of reach and grasp in quadriplegic patients. We are addressing the question of how to model neuromusculoskeletal systems so that they are sufficiently complete, valid and accessible to be useful in both basic and applied sensorimotor research.

Biomechanical Phenomena↗

Integration of genomic datasets to predict protein complexes in yeast.

The ultimate goal of functional genomics is to define the function of all the genes in the genome of an organism. A large body of information of the biological roles of genes has been accumulated and aggregated in the past decades of research, both from traditional experiments detailing the role of individual genes and proteins, and from newer experimental strategies that aim to characterize gene function on a genomic scale. It is clear that the goal of functional genomics can only be achieved by integrating information and data sources from the variety of these different experiments. Integration of different data is thus an important challenge for bioinformatics. The integration of different data sources often helps to uncover non-obvious relationships between genes, but there are also two further benefits. First, it is likely that whenever information from multiple independent sources agrees, it should be more valid and reliable. Secondly, by looking at the union of multiple sources, one can cover larger parts of the genome. This is obvious for integrating results from multiple single gene or protein experiments, but also necessary for many of the results from genome-wide experiments since they are often confined to certain (although sizable) subsets of the genome. In this paper, we explore an example of such a data integration procedure. We focus on the prediction of membership in protein complexes for individual genes. For this, we recruit six different data sources that include expression profiles, interaction data, essentiality and localization information. Each of these data sources individually contains some weakly predictive information with respect to protein complexes, but we show how this prediction can be improved by combining all of them. Supplementary information is available at http:// bioinfo.mbb.yale.edu/integrate/interactions/.

Cell Cycle↗

An integrated approach for finding overlooked genes in yeast.

We report here the discovery of 137 previously unappreciated genes in yeast through a widely applicable and highly scalable approach integrating methods of gene-trapping, microarray-based expression analysis, and genome-wide homology searching. Our approach is a multistep process in which expressed sequences are first trapped using a modified transposon that produces protein fusions to beta-galactosidase (beta-gal); non-annotated open reading frames (ORFs) translated as beta-gal chimeras are selected as a candidate pool of potential genes. To verify expression of these sequences, labeled RNA is hybridized against a microarray of oligonucleotides designed to detect gene transcripts in a strand-specific manner. In complement to this experimental method, novel genes are also identified in silico by homology to previously annotated proteins. As these methods are capable of identifying both short ORFs and antisense ORFs, our approach provides an effective supplement to current gene-finding schemes. In total, the genes discovered using this approach constitute 2% of the yeast genome and represent a wealth of overlooked biology.

Chromosomes↗

Stability analysis for postural control in a two-joint limb system.

The stability behavior of a multi-joint limb with electrically activated muscles provides important clues for postural control of motor tasks. The stability property of the musculoskeletal system can be characterized with its eigenvalues evaluated at operating postures in the workspace. A planar arm model with shoulder and elbow joints and three pairs of antagonistic muscles was constructed in ADAMS. Stability behavior of shoulder and elbow joints was analyzed using the loci of eigenvalues in the s-plane. In the analysis of open-loop cocontraction of antagonist muscles with increasing activation from 5% to 100%, the eigenvalues of the shoulder and elbow joints were confined within the left half of the s-plane in a stripe of +/- j 0.5, and moved toward left onto the real axis. The shoulder eigenvalues were generally nearer to the imaginary axis than the elbow ones, indicating a more oscillatory behavior at the shoulder joint than that at the elbow joint. The effects of joint configuration evaluated within the workspace from 40 degrees to 110 degrees for the elbow and from 40 degrees to 120 degrees for the shoulder showed that the elbow eigenvalues were more prone to configuration changes, particularly elbow angles. We also developed a simulation paradigm for sampled data FES control systems that contain a mixture of continuous time components and sampling and hold effects. This simulation paradigm is useful for realistic simulation of local feedback controller performance.

Computer Simulation↗