Search PubMed⌕ Search

Biomedical subjects

L Allison

Publications and source records attributed to L Allison.

At least 19 recordsLinked to original sources

Discovering patterns in Plasmodium falciparum genomic DNA.

A method has been developed for discovering patterns in DNA sequences. Loosely based on the well-known Lempel Ziv model for text compression, the model detects repeated sequences in DNA. The repeats can be forward or inverted, and they need not be exact. The method is particularly useful for detecting distantly related sequences, and for finding patterns in sequences of biased nucleotide composition, where spurious patterns are often observed because the bias leads to coincidental nucleotide matches. We show here the utility of the method by applying it to genomic sequences of Plasmodium falciparum. A single scan of chromosomes 2 and 3 of P. falciparum, using our method and no other a priori information about the sequences, reveals regions of low complexity in both telomeric and central regions, long repeats in the subtelomeric regions, and shorter repeat areas in dense coding regions. Application of the method to a recently sequenced contig of chromosome 10 that has a particularly biased base composition detects a long internal repeat more readily than does the conventional dot matrix plot. Space requirements are linear, so the method can be used on large sequences. The observed repeat patterns may be related to large-scale chromosomal organization and control of gene expression. The method has general application in detecting patterns of potential interest in newly sequenced genomic material.

Algorithms↗

Fast, optimal alignment of three sequences using linear gap costs.

Alignment algorithms can be used to infer a relationship between sequences when the true relationship is unknown. Simple alignment algorithms use a cost function that gives a fixed cost to each possible point mutation-mismatch, deletion, insertion. These algorithms tend to find optimal alignments that have many small gaps. It is more biologically plausible to have fewer longer gaps rather than many small gaps in an alignment. To address this issue, linear gap cost algorithms are in common use for aligning biological sequence data. More reliable inferences are obtained by aligning more than two sequences at a time. The obvious dynamic programming algorithm for optimally aligning k sequences of length n runs in O(n(k)) time. This is impractical if k>/=3 and n is of any reasonable length. Thus, for this problem there are many heuristics for aligning k sequences, however, they are not guaranteed to find an optimal alignment. In this paper, we present a new algorithm guaranteed to find the optimal alignment for three sequences using linear gap costs. This gives the same results as the dynamic programming algorithm for three sequences, but typically does so much more quickly. It is particularly fast when the (three-way) edit distance is small. Our algorithm uses a speed-up technique based on Ukkonen's greedy algorithm (Ukkonen, 1983) which he presented for two sequences and simple costs.

Algorithms↗

Sequence complexity for biological sequence analysis.

A new statistical model for DNA considers a sequence to be a mixture of regions with little structure and regions that are approximate repeats of other subsequences, i.e. instances of repeats do not need to match each other exactly. Both forward- and reverse-complementary repeats are allowed. The model has a small number of parameters which are fitted to the data. In general there are many explanations for a given sequence and how to compute the total probability of the data given the model is shown. Computer algorithms are described for these tasks. The model can be used to compute the information content of a sequence, either in total or base by base. This amounts to looking at sequences from a data-compression point of view and it is argued that this is a good way to tackle intelligent sequence analysis in general.

Algorithms↗

Clinical effectiveness: testicular self-examination.

Evidence must be used to support a change in practice. Incorporating evidence into practice is essential for clinical effectiveness. The findings should be disseminated to others involved. The effect on patients should be evaluated.

Health Education↗

Genetic heterogeneity of Escherichia coli O157:H7 in Scotland and its utility in strain subtyping.

From April 1994 to March 1995, seven outbreaks of Escherichia coli O157:H7 infection occurred throughout Scotland, including the largest milk-borne outbreak to date worldwide. Various vehicles of infection were identified, and there were 144 confirmed cases in total. All isolates associated with the outbreaks were subjected to detailed subtyping: phage typing, testing for carriage of verotoxin genes (VT), and pulsed-field gel electrophoresis. The outbreak strains were of three different phage types (2, 4, and 28). Those of phage type 2 and 28 were VT1-/VT2+, those of phage type 4 were VT1+/VT2+. To discriminate outbreak-associate isolates from the high sporadic background, real-time pulsed-field gel electrophoresis analyses were performed. The results demonstrated that, within each of the seven outbreak groups, the macrorestriction profiles observed were indistinguishable, whereas profiles for sporadic isolates were not. The consistent genetic heterogeneity observed within the Scottish Escherichia coli O157 population can be exploited in epidemiological investigations.

Animals↗

rbcL Transcript levels in tobacco plastids are independent of light: reduced dark transcription rate is compensated by increased mRNA stability.

The plastid rbcL gene, encoding the large subunit of ribulose-1, 5-bisphosphate carboxylase, in higher plants is transcribed from a sigma70 promoter by the eubacterial-type RNA polymerase. To identify regulatory elements outside of the rbcL -10/-35 promoter core, we constructed transplastomic tobacco plants with uidA reporter genes expressed from rbcL promoter derivatives. Promoter activity was characterized by measuring steady state levels of uidA mRNA on RNA gel blots and by measuring promoter strength in run-on transcription assays. We report here that the rbcL core promoter is sufficient to obtain wild-type rates of transcription. Furthermore, the rates of transcription were up to 10-fold higher in light-grown leaves than in dark-adapted plants. Although the rates of transcription were lower in the dark, rbcL mRNA accumulated to similar levels in light-grown and dark-adapted leaves. Accumulation of uidA mRNA from most rbcL promoter deletion derivatives directly reflected the relative rates of transcription: high in the light-grown and low in the dark-adapted leaves. However, uidA mRNA accumulated to high levels in a light-independent fashion as long as a segment encoding a stem-loop structure in the 5' untranslated region was included in the promoter construct. This finding indicates that lower rates of rbcL transcription in the dark are compensated by increased mRNA stability.

Base Sequence↗

An MML classification of protein structure that knows about angles and sequence.

The MML classification program, Snob, deals with mixture modelling (or clustering) of circular data. It has recently been extended to do Markov modelling of the serial correlation between clusters such as modelling the fact that a Helix cluster favours being followed by another Helix cluster. Such a model is better known as a Hidden Markov Model. The search for the most appropriate secondary structure classification of protein data is of significant importance and was addressed by Hunter and States (1992) using the Bayesian classifier, AutoClass, on Cartesian co-ordinate data of protein residues. Dowe et al. (1996) improved upon this earlier work by using Snob to cluster dihedral angle data, with the advantage that 3 x 3 = 9 Cartesian co-ordinates can be represented by the 2 orientation-invariant angles, phi and psi. The Hidden Markov Model used here is shown to be a more appropriate way again of modelling protein data and results in the selection of a simpler class model with 17 structure classes. We report on this classification, including the class transition matrix, and relate it back to the amino-acid sequence and the simple Helix, Beta, Turn classification. We find 3 types of Helix, 2 types of Beta and many types of Turn. The msot numerous Turn class defines a continuous flexible structure that is negatively correlated to all the other classes.

Amino Acid Sequence↗

Discovering simple DNA sequences by compression.

An information-theoretic DNA compression scheme devised by Milosavljevic and Jurka (1993) has been used in many places in the literature for both the discovery of new genes and the compression of DNA. Their compression method applies an encoding of previously occurring runs. They use 5 different code-words: four being the DNA bases, A, C, G and T, and the other being a pointer to a previously occurring run. They advocate a code-word of length log2 5 for each of these and then encoding a run by a code-word of length 2 x log2 n, where n is the length of the sequence. This scheme encodes the start of the sequence with a code-word of length log2 n and likewise encodes the end of the sequence with a code-word of length log2 n. In this paper, we show the above coding scheme to be inefficient in various ways and improve upon it so that it can compress DNA. We discuss our implementation of various schemes some of which run in linear time.

Algorithms↗

Compression of strings with approximate repeats.

We describe a model for strings of characters that is loosely based on the Lempel Ziv model with the addition that a repeated substring can be an approximate match to the original substring; this is close to the situation of DNA, for example. Typically there are many explanations for a given string under the model, some optimal and many suboptimal. Rather than commit to one optimal explanation, we sum the probabilities over all explanations under the model because this gives the probability of the data under the model. The model has a small number of parameters and these can be estimated from the given string by an expectation-maximization (EM) algorithm. Each iteration of the EM algorithm takes O(n2) time and a few iterations are typically sufficient. O(n2) complexity is impractical for strings of more than a few tens of thousands of characters and a faster approximation algorithm is also given. The model is further extended to include approximate reverse complementary repeats when analyzing DNA strings. Tests include the recovery of parameter estimates from known sources and applications to real DNA strings.

Algorithms↗

Relationship between head growth and neurodevelopmental outcome of Malaysian very low birthweight infants during the 1st year of life.

A prospective study was carried out to (i) compare head growth patterns of 103 very low birthweight (VLBW, < 1500 g) Malaysian infants and 98 normal birthweight (NBW, 2500- < 4500 g) controls during the 1st year of life; and (ii) examine the relationship between neurodevelopmental outcome at 1 year of age and occipito-frontal head circumferences (OFC) at birth and at 1 year of age in VLBW babies. When compared with those of NBW infants at birth, mid-infancy and 1 year of age, the mean OFC ratios (observed/expected OFC at 50th percentile) of VLBW infants were significantly lower (p < 0.001). Small-for-gestational-age (SGA) VLBW babies had significantly lower mean OFC ratios than their appropriate-for-gestational-age (AGA) VLBW counterparts at birth (p < 0.001), but this difference was no longer seen at mid-infancy or at 1 year of age. Logistic regression analysis showed that abnormal late neonatal cranial ultrasound findings (odds ratio 8.5, 95% confidence interval 4.12-22.07; p < 0.001) and each additional day of oxygen therapy (odds ratio 1.15, 95% confidence interval 1.00-4.45; p = 0.045) were significant risk factors associated with neurodevelopmental disability at 1 year of age, while mean OFC ratios at birth or at 1 year of age were not. Poor postnatal head growth per se did not predict disability, but probably reflected the consequences of "brain injury" as evidenced by abnormal brain scans.

Adult↗

Neurodevelopmental outcome of Malaysian very low birth weight infants: predictive value of cranial ultrasound appearances.

The aim of the study was to determine the predictive value of cranial ultrasound scans done in the neonatal period for neurodevelopmental outcome of the Malaysian very low birthweight (VLBW, < 1500 grams) infants assessed at 12 months of corrected age. Of the 101 infants studied, 68 (67.3%) were neurodevelopmentally normal at one year of age, 18 (17.8%) had major and 15 (14.9%) had minor neurodevelopmental impairment. Neurodevelopmental outcome was normal in 66/88 (75.0%) infants who did not have severe intraventricular haemorrhage (IVH) or periventricular intraparenchymal echo densities (PVE) in the first week of life, and in 57/73 (78.1%) with uncomplicated scans at discharge. In contrast, 11/13 (84.6%) with parenchymal echo densities or severe intraventricular bleed in the early neonatal period and 17/28 (60.7%) with complicated scans at discharge had adverse sequelae. There was a significant association between lesions seen on cranial ultrasound in the neonatal period and subsequent neurodevelopmental impairment. Late neonatal ultrasound scans appear to be a better predictor of short-term neurodevelopmental outcome than early scans.

Child Development↗

Comparison of morbidities in very low birthweight and normal birthweight infants during the first year of life in a developing country.

OBJECTIVE: To compare the morbidities in the very low birthweight (VLBW; < 1500 g) and normal birthweight (NBW; > or = 2500 g) Malaysian infants during the first year of life. METHODOLOGY: Prospective observational cohort study of consecutive surviving VLBW infants and randomly sampled NBW infants born in the Kuala Lumpur Maternity Hospital between 1 December 1989 and 31 December 1992. Infants were followed up regularly during the first year of life, after correction for prematurity. RESULTS: Compared with NBW infants (n = 106), VLBW infants (n = 127) had significantly higher risk of failure to thrive (odds ratio [OR] = 8.0, 95% confidence intervals [CI]: 1.1 to 354.3), wheezing (OR = 3.7, 95% CI: 1.6 to 9.3), rehospitalization (OR = 2.3, 95% CI: 1.1 to 5.0), cerebral palsy (OR = 8.6, 95% CI: 2.0 to 77.6), neurosensory hearing loss (OR = 12.0, 95% CI: 1.7 to 513.6) and visual loss (7.9 vs 0%, P = 0.002). The mean mental developmental index (MDI) and mean psychomotor developmental index (PDI) at 1 year of age were significantly lower among VLBW infants (MDI 99 [SD = 28], PDI 89 [SD = 25]) than NBW infants (MDI 106 [SD = 18], PDI 101 [SD = 18]) (95% CI for difference between means being MDI: -14.1 to -1.7; and PDI: -17.6 to -6.0). Logistic regression analysis showed that among VLBW infants: (i) male sex, Malay ethnicity and bronchopulmonary dysplasia were significant risk factors associated with wheezing; (ii) longer duration of oxygen therapy during the neonatal period, seizures after the post-neonatal period and wheezing were significant risk factors associated with rehospitalization; and (iii) longer duration of oxygen therapy during the neonatal period was a significant risk factor associated with adverse neurodevelopmental outcome during the first year of life. CONCLUSIONS: Compared with NBW infants, VLBW Malaysian infants had significantly higher risks of physical and neuro-developmental morbidities.

Case-Control Studies↗

Circular clustering of protein dihedral angles by Minimum Message Length.

Early work on proteins identified the existence of helices and extended sheets in protein secondary structures, a high-level classification which remains popular today. Using the Snob program for information-theoretic Minimum Message Length (MML) classification, we are able to take the protein dihedral angles as determined by X-ray crystallography, and cluster sets of dihedral angles into groups. Previous work by Hunter and States has applied a similar Bayesian classification method, AutoClass, to protein data with site position represented by 3 Cartesian co-ordinates for each of the alpha-Carbon, beta-Carbon and Nitrogen, totalling 9 co-ordinates. By using the von Mises circular distribution in the Snob program, we are instead able to represent local site properties by the two dihedral angles, phi and psi. Since each site can be modelled as having 2 degrees of freedom, this orientation-invariant dihedral angle representation of the data is more compact than that of nine highly-correlated Cartesian co-ordinates. Using the information-theoretic message length concepts discussed in the paper, such a more concise model is more likely to represent the underlying generating process from which the data came. We report on the results of our classification, plotting the classes in (phi, psi) space; and introducing a symmetric information-theoretic distance measure to build a minimum spanning tree between the classes. We also give a transition matrix between the classes and note the existence of three classes in the region phi approximately -1.09 rad and psi approximately -0.75 rad which are close on the spanning tree and have high inter-transition probabilities. This gives rise to a tight, abundant and self-perpetuating structure.

Computational Biology↗

Long-term efficacy of continuing hepatitis B vaccination in infancy in two Gambian villages.

In 1984, all non-immune children under the age of 5 years in the Gambian villages of Keneba and Manduar were vaccinated against hepatitis B virus (HBV). All children born in these villages since 1984 have been vaccinated in infancy. Despite a rapid fall in antibody concentrations, vaccine efficacy against HBV infection and chronic carriage of HBsAg has increased with time. Overall, vaccine efficacies in 1993 against HBV infection and chronic HBsAg carriage were 94.7% (95% Cl 93.0-96.0) and 95.3% (91.0-97.5), respectively. Breakthrough infections in vaccinated children largely originate from chronic HBsAg carriers. Thus, we tested 261 chronic carriers for HBV DNA and e antigen. The prevalence of these markers of infectivity, and the amount of HBV DNA, decreased greatly with age. Detailed studies of breakthrough infections over two 4-year periods revealed that in the second period there were fewer than half the expected numbers of infections. Our findings suggest that in Keneba and Manduar long-term vaccination is progressively decreasing HBV transmission by chronic carriers, since their infectivity diminishes with time.

Adolescent↗

The posterior probability distribution of alignments and its application to parameter estimation of evolutionary trees and to optimization of multiple alignments.

How to sample alignments from their posterior probability distribution given two strings is shown. This is extended to sampling alignments of more than two strings. The result is first applied to the estimation of the edges of a given evolutionary tree over several strings. Second, when used in conjunction with simulated annealing, it gives a stochastic search method for an optimal multiple alignment.

Algorithms↗

Breakthrough infections and identification of a viral variant in Gambian children immunized with hepatitis B vaccine.

Hepatitis B (HB) breakthrough infections, identified by the presence of HB core (c) antibody, were found in 32 of 358 Gambian children vaccinated with plasma-derived HB vaccine. Over 2 years, 15 of these children lost their HBc antibodies. These children had significantly higher HB surface antibody levels before infection than those who retained HBc antibodies. One child, who responded well to the vaccine, had HB viral DNA detected in the presence of HBs antibodies. The S gene sequence of this DNA showed nucleotide changes that resulted in an amino acid substitution at residue 141 (lysine to glutamic acid) of the surface antigen. This finding suggests the child was infected with a variant virus that was not neutralized by antibodies resulting from HB vaccination.

Child↗