Search PubMed⌕ Search

Biomedical subjects

L Hillier

Publications and source records attributed to L Hillier.

At least 19 recordsLinked to original sources

zA map for sequence analysis of the Arabidopsis thaliana genome.

Arabidopsis thaliana has emerged as a model system for studies of plant genetics and development, and its genome has been targeted for sequencing by an international consortium (the Arabidopsis Genome Initiative; http://genome-www. stanford.edu/Arabidopsis/agi.html). To support the genome-sequencing effort, we fingerprinted more than 20,000 BACs (ref. 2) from two high-quality publicly available libraries, generating an estimated 17-fold redundant coverage of the genome, and used the fingerprints to nucleate assembly of the data by computer. Subsequent manual revision of the assemblies resulted in the incorporation of 19,661 fingerprinted BACs into 169 ordered sets of overlapping clones ('contigs'), each containing at least 3 clones. These contigs are ideal for parallel selection of BACs for large-scale sequencing and have supported the generation of more than 5.8 Mb of finished genome sequence submitted to GenBank; analysis of the sequence has confirmed the integrity of contigs constructed using this fingerprint data. Placement of contigs onto chromosomes can now be performed, and is being pursued by groups involved in both sequencing and positional cloning studies. To our knowledge, these data provide the first example of whole-genome random BAC fingerprint analysis of a eucaryote, and have provided a model essential to efforts aimed at generating similar databases of fingerprint contigs to support sequencing of other complex genomes, including that of human.

Arabidopsis↗

An encyclopedia of mouse genes.

The laboratory mouse is the premier model system for studies of mammalian development due to the powerful classical genetic analysis possible (see also the Jackson Laboratory web site, http://www.jax.org/) and the ever-expanding collection of molecular tools. To enhance the utility of the mouse system, we initiated a program to generate a large database of expressed sequence tags (ESTs) that can provide rapid access to genes. Of particular significance was the possibility that cDNA libraries could be prepared from very early stages of development, a situation unrealized in human EST projects. We report here the development of a comprehensive database of ESTs for the mouse. The project, initiated in March 1996, has focused on 5' end sequences from directionally cloned, oligo-dT primed cDNA libraries. As of 23 October 1998, 352,040 sequences had been generated, annotated and deposited in dbEST, where they comprised 93% of the total ESTs available for mouse. EST data are versatile and have been applied to gene identification, comparative sequence analysis, comparative gene mapping and candidate disease gene identification, genome sequence annotation, microarray development and the development of gene-based map resources.

Animals↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

DNA sequence chromatogram browsing using JAVA and CORBA.

DNA sequence chromatograms (traces) are the primary data source for all large-scale genomic and expressed sequence tags (ESTs) sequencing projects. Access to the sequencing trace assists many later analyses, for example contig assembly and polymorphism detection, but obtaining and using traces is problematic. Traces are not collected and published centrally, they are much larger than the base calls derived from them, and viewing them requires the interactivity of a local graphical client with local data. To provide efficient global access to DNA traces, we developed a client/server system based on flexible Java components integrated into other applications including an applet for use in a WWW browser and a stand-alone trace viewer. Client/server interaction is facilitated by CORBA middleware which provides a well-defined interface, a naming service, and location independence. [The software is packaged as a Jar file available from the following URL: http://www.ebi.ac.uk/jparsons. Links to working examples of the trace viewers can be found at http://corba.ebi.ac.uk/EST. All the Washington University mouse EST traces are available for browsing at the same URL.]

Animals↗

Single nucleotide polymorphism hunting in cyberspace.

Large-scale sequencing of human cDNA and genomic DNA libraries has produced a large collection of sequence data in public databases. To date, >900,000 human expressed sequence tag (EST) sequences and >80,000,000 bases of genomic DNA sequence have been deposited in Genbank. This ever-expanding data set is a rich source of gene-associated and anonymous single nucleotide polymorphisms (SNPs). DNA sequence variations can be found by comparing the sequences of redundant ESTs and by comparing sequences from overlapping genomic clones. Initial studies have shown that, with proper computer screening, informative SNP markers can be developed from these DNA databases in an efficient and cost-effective manner. Complete public access to these databases will allow individual investigators to add biological value to the human sequence data generated by large-scale sequencing centers.

DNA↗

"When you carry condoms all the boys think you want it": negotiating competing discourses about safe sex.

With the advent of HIV, sexual health campaigns and formal sex education in schools have worked to instil the concept of safe sex into the collective minds of Australia's youth. However the concept in its present guise is a fairly limited one. We argue in this paper that the predominant emphasis in education programmes on safe sex as condom use may be counter-productive for some young heterosexuals for two reasons. First, this strategy is male-focused and may not extrapolate well to young women who face special risks around pregnancy and rigid societal gender norms which govern sexual behaviour. Second, health promotion strategies aimed at young heterosexuals are based on an assumption of rational decision-making in sexual encounters and obscure the non-rational nature of arousal and desire, and the unequal power relations that exist between young men and women engaging in sex. Five hundred and twelve senior rural students participated in the study which included group discussions about sexuality and survey items which focused on the meanings of safe sex and the accessibility and use of condoms. The results showed that though most students identified condoms with safe sex, many were ambivalent about using them. Reasons given related to problems of negotiation, difficulties of access, and the risks which condoms gave no protection from, such as a sullied reputation. Perhaps, partly because of this, some students were looking to less secure methods of protection such as informal history-taking and monogamy. It is argued that successful sexual health promotion strategies must address the broad spectrum of concerns facing young men and women when they become sexually active and that consideration be given to the social context in which young people conduct their sexual lives.

Adolescent↗

Gene discovery by EST sequencing in Toxoplasma gondii reveals sequences restricted to the Apicomplexa.

To accelerate gene discovery and facilitate genetic mapping in the protozoan parasite Toxoplasma gondii, we have generated >7000 new ESTs from the 5' ends of randomly selected tachyzoite cDNAs. Comparison of the ESTs with the existing gene databases identified possible functions for more than 500 new T. gondii genes by virtue of sequence motifs shared with conserved protein families, including factors involved in transcription, translation, protein secretion, signal transduction, cytoskeleton organization, and metabolism. Despite this success in identifying new genes, more than 50% of the ESTs correspond to genes of unknown function, reflecting the divergent evolutionary status of this parasite. A newly recognized class of genes was identified based on its similarity to sequences known only from other members of the same phylum, therefore identifying sequences that are apparently restricted to the Apicomplexa. Such genes may underlie pathways common to this group of medically important parasites, therefore identifying potential targets for intervention.

Animals↗

Base-calling of automated sequencer traces using phred. I. Accuracy assessment.

The availability of massive amounts of DNA sequence information has begun to revolutionize the practice of biology. As a result, current large-scale sequencing output, while impressive, is not adequate to keep pace with growing demand and, in particular, is far short of what will be required to obtain the 3-billion-base human genome sequence by the target date of 2005. To reach this goal, improved automation will be essential, and it is particularly important that human involvement in sequence data processing be significantly reduced or eliminated. Progress in this respect will require both improved accuracy of the data processing software and reliable accuracy measures to reduce the need for human involvement in error correction and make human review more efficient. Here, we describe one step toward that goal: a base-calling program for automated sequencer traces, phred, with improved accuracy. phred appears to be the first base-calling program to achieve a lower error rate than the ABI software, averaging 40%-50% fewer errors in the data sets examined independent of position in read, machine running conditions, or sequencing chemistry.

Algorithms↗

Sequence assembly with CAFTOOLS.

Large-scale genomic sequencing requires a software infrastructure to support and integrate applications that are not directly compatible. We describe a suite of software tools built around the Common Assembly Format (CAF), a comprehensive representation of a sequence assembly as a text file. These tools form the backbone of sequencing informatics at the Sanger Centre and the Genome Sequencing Center. The CAF format is intentionally flexible, and our Perl and C libraries, which parse and manipulate it, provide powerful tools for creating new applications as well as wrappers to incorporate other software. The tools are available free by anonymous FTP from ftp://ftp.sanger.ac.uk/pub/badger/.

Algorithms↗

Overlapping genomic sequences: a treasure trove of single-nucleotide polymorphisms.

An efficient strategy to develop a dense set of single-nucleotide polymorphism (SNP) markers is to take advantage of the human genome sequencing effort currently under way. Our approach is based on the fact that bacterial artificial chromosomes (BACs) and P1-based artificial chromosomes (PACs) used in long-range sequencing projects come from diploid libraries. If the overlapping clones sequenced are from different lineages, one is comparing the sequences from 2 homologous chromosomes in the overlapping region. We have analyzed in detail every SNP identified while sequencing three sets of overlapping clones found on chromosome 5p15.2, 7q21-7q22, and 13q12-13q13. In the 200.6 kb of DNA sequence analyzed in these overlaps, 153 SNPs were identified. Computer analysis for repetitive elements and suitability for STS development yielded 44 STSs containing 68 SNPs for further study. All 68 SNPs were confirmed to be present in at least one of the three (Caucasian, African-American, Hispanic) populations studied. Furthermore, 42 of the SNPs tested (62%) were informative in at least one population, 32 (47%) were informative in two or more populations, and 23 (34%) were informative in all three populations. These results clearly indicate that developing SNP markers from overlapping genomic sequence is highly efficient and cost effective, requiring only the two simple steps of developing STSs around the known SNPs and characterizing them in the appropriate populations.

Bacteriophage P1↗

Automated sequence preprocessing in a large-scale sequencing environment.

A software system for transforming fragments from four-color fluorescence-based gel electrophoresis experiments into assembled sequence is described. It has been developed for large-scale processing of all trace data, including shotgun and finishing reads, regardless of clone origin. Design considerations are discussed in detail, as are programming implementation and graphic tools. The importance of input validation, record tracking, and use of base quality values is emphasized. Several quality analysis metrics are proposed and applied to sample results from recently sequenced clones. Such quantities prove to be a valuable aid in evaluating modifications of sequencing protocol. The system is in full production use at both the Genome Sequencing Center and the Sanger Centre, for which combined weekly production is approximately 100, 000 sequencing reads per week.

Automation↗

Rural youth: HIV/STD knowledge levels and sources of information.

Though it is recognised that a sound knowledge of HIV/STD transmission is insufficient on its own to ensure that young people practise safer sex, knowledge of this sort is a necessary prerequisite to safe sex behaviours. Research on the HIV/STD knowledge levels of young people has shown that, while they have high levels of HIV knowledge, they generally have little idea about the transmission of other STDs. Moreover, most of this research has been conducted with more accessible youth in large urban centres. The research reported in this paper focused on HIV/STD knowledge levels, awareness of safe and unsafe sexual practices, and the information sources accessed by 1168 young people living in small rural towns in Queensland, Tasmania and Victoria. As with mainstream youth, the participants in this study had high levels of HIV knowledge and very low levels of knowledge about the transmission of other STDs, their names and symptoms. Practical knowledge of the safety of sexual practices was good, although a number of students confused high risk practices with high risk groups. Students mainly accessed informal information sources such as parents, peers and magazines and took little advantage of more formal sources even when they were available in their towns. There were few gender differences in levels of knowledge; however, girls accessed more information sources and knew more names of STDs. The implications of these findings are discussed in relation to young people developing critical attitudes towards more informal sources, and the better use of more formal sources.

Adolescent↗

Expressed sequence tag analysis of the bradyzoite stage of Toxoplasma gondii: identification of developmentally regulated genes.

Toxoplasma gondii is a protozoan parasite responsible for widespread infections in humans and animals. Two major asexual forms are produced during the life cycle of this parasite: the rapidly dividing tachyzoite and the more slowly dividing, encysted bradyzoite. To further study the differentiation between these two forms, we have generated a large number of expressed sequence tags (ESTs) from both asexual stages. Previously, we obtained data on approximately 7,400 ESTs from tachyzoites (J. Ajioka et al., Genome Res. 8:18-28, 1998). Here, we report the results from analysis of approximately 2,500 ESTs from bradyzoites purified from the cysts of infected mice. We also report the results from analysis of 760 ESTs from parasites induced to differentiate from tachyzoites to bradyzoites in vitro. Comparison of the data sets from bradyzoites and tachyzoites reveals many previously uncharacterized sequence clusters which are largely or completely specific to one or other developmental stage. This class includes a bradyzoite-specific form of enolase. Combined with the previously identified bradyzoite-specific form of lactate dehydrogenase, this finding suggests significant differences in flux through the lower end of the glycolytic pathway in this stage. Thus, the generation of this data set provides valuable insights into the metabolism and growth of the parasite in the encysted form and represents a substantial body of information for further study of development in Toxoplasma.

Amino Acid Sequence↗

The homozygous complete hydatidiform mole: a unique resource for genome studies.

The most frequent type of complete hydatidiform mole is a 46, XX homozygote formed by the fertilization of an empty ovum by a single haploid sperm that later duplicates its chromosomes to give a diploid tumor. The homozygous nature of these complete hydatidiform moles makes them unique resources for human genome studies. They can serve as homozygous controls in the development of single nucleotide polymorphism (SNP) markers and provide a way to obtain long-range haplotypes that are useful in population studies. The use of a homozygous control makes it possible to estimate the allele frequencies of the SNP markers in any population by sequencing pooled DNA samples. In this report, we present evidence of homozygosity of a complete hydatidiform mole using 20 diallelic markers distributed across the genome. Furthermore, its usefulness as a homozygous control in SNP development and as a resource for long-range haplotype determination is demonstrated using 11 newly discovered loci in the BRCA2 region on chromosome 13q12-q13.

BRCA2 Protein↗

Representation of cloned genomic sequences in two sequencing vectors: correlation of DNA sequence and subclone distribution.

Representation of subcloned Caenorhabditis elegans and human DNA sequences in both M13 and pUC sequencing vectors was determined in the context of large scale genomic sequencing. In many cases, regions of subclone under-representation correlated with the occurrence of repeat sequences, and in some cases the under-representation was orientation specific. Factors which affected subclone representation included the nature and complexity of the repeat sequence, as well as the length of the repeat region. In some but not all cases, notable differences between the M13 and pUC subclone distributions existed. However, in all regions lacking one type of subclone (either M13 or pUC), an alternate subclone was identified in at least one orientation. This suggests that complementary use of M13 and pUC subclones would provide the most comprehensive subclone coverage of a given genomic sequence.

Animals↗

The nucleotide sequence of Saccharomyces cerevisiae chromosome XII.

The yeast Saccharomyces cerevisiae is the pre-eminent organism for the study of basic functions of eukaryotic cells. All of the genes of this simple eukaryotic cell have recently been revealed by an international collaborative effort to determine the complete DNA sequence of its nuclear genome. Here we describe some of the features of chromosome XII.

Base Sequence↗

'That's the problem with living in a small town': privacy and sexual health issues for young rural people.

Survey and focus group discussions examining sexual health issues for young people were conducted with 1168 year 8 and year 10 secondary school students living in small rural communities across Australia. Growing up in the country was generally perceived as a positive experience; however, many young people felt that they had little privacy. Two main areas of concern emerged in relation to sexual health issues: worries about being recognised in public venues such as doctor's surgeries and chemists, and the informal mechanisms among peer groups that appraised and regulated sexual behaviour and attitudes. There were some significant gender differences evident in the expectations and experiences attached to these concerns, with girls expressing more awareness of and concern towards their public reputations when accessing sexual health services. They also felt that their sexual reputations among peers were closely monitored by way of their behaviour and appearance. This can militate against confident and assertive safer sex strategies such as condom use, when initiating or insisting upon condom use is construed as evidence of promiscuity or a preparedness to engage in sex. Concerns around privacy may be acutely experienced by young rural women, and health services providers need to be aware of these issues and efforts need to be made to address and allay the apprehensions of young people in this sensitive area.

Adolescent↗