Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Pattern Recognition, Automated”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Robust quantification of in vitro angiogenesis through image analysis.

An automated image analysis method for quantification of in vitro angiogenesis is presented. The method is designed for in vitro angiogenesis assays that are based on co-culturing endothelial cells with fibroblasts. Such assays are used in many current studies in which anti-angiogenic agents for the treatment of cancer are being sought. This search requires accurate quantification of the stimulatory and inhibitory effects of the different agents. The quantification method gives lengths and sizes of the tubule complexes as well as the numbers of junctions in each of them. The method is tested with a set of test images obtained with a commercially available in vitro angiogenesis assay. The results correctly indicate the inhibitory effect of suramin and the stimulatory effect of vascular endothelial growth factor. Moreover, the image analysis method is shown to be robust against variations in illumination. We have implemented a software package that utilizes the methods. The software as well as a set of test images are available at http://www.cs.tut.fi/sgn/csb/angioquant/.

Algorithms↗

Automated evaluation of Her-2/neu status in breast tissue from fluorescent in situ hybridization images.

The evaluation of fluorescent in situ hybridization (FISH) images is one of the most widely used methods to determine Her-2/neu status of breast samples, a valuable prognostic indicator. Conventional evaluation is a difficult task since it involves manual counting of dots in multiple images. In this paper, we present a multistage algorithm for the automated classification of FISH images from breast carcinomas. The algorithm focuses not only on the detection of FISH dots per image, but also on combining results from multiple images taken from a slice for overall case classification. The algorithm includes mainly two stages for nuclei and dot detection respectively. The dot segmentation consists of a top-hat filtering stage followed by template matching to separate real signals from noise. Nuclei segmentation includes a nonlinearity correction step, global thresholding to identify candidate regions, and a geometric rule to distinguish between holes within a nucleus and holes between nuclei. Finally, the marked watershed transform is used to segment cell nuclei with markers detected as regional maxima of the distance transform. Combining the two stages allows the measurement of FISH signals ratio per cell nucleus and the collective classification of cases as positive or negative. The system was evaluated with receiver operating characteristic analysis and the results were encouraging for the further development of this method.

Algorithms↗

Adaptive elastic segmentation of brain MRI via shape-model-guided evolutionary programming.

This paper presents a fully automated segmentation method for medical images. The goal is to localize and parameterize a variety of types of structure in these images for subsequent quantitative analysis. We propose a new hybrid strategy that combines a general elastic template matching approach and an evolutionary heuristic. The evolutionary algorithm uses prior statistical information about the shape of the target structure to control the behavior of a number of deformable templates. Each template, modeled in the form of a B-spline, is warped in a potential field which is itself dynamically adapted. Such a hybrid scheme proves to be promising: by maintaining a population of templates, we cover a large domain of the solution space under the global guidance of the evolutionary heuristic, and thoroughly explore interesting areas. We address key issues of automated image segmentation systems. The potential fields are initially designed based on the spatial features of the edges in the input image, and are subjected to spatially adaptive diffusion to guarantee the deformation of the template. This also improves its global consistency and convergence speed. The deformation algorithm can modify the internal structure of the templates to allow a better match. We investigate in detail the preprocessing phase that the images undergo before they can be used more effectively in the iterative elastic matching procedure: a texture classifier, trained via linear discriminant analysis of a learning set, is used to enhance the contrast of the target structure with respect to surrounding tissues. We show how these techniques interact within a statistically driven evolutionary scheme to achieve a better tradeoff between template flexibility and sensitivity to noise and outliers. We focus on understanding the features of template matching that are most beneficial in terms of the achieved match. Examples from simulated and real image data are discussed, with considerations of algorithmic efficiency.

Algorithms↗

A Monte Carlo approach to rolling leukocyte tracking in vivo.

Tracking the movement of rolling leukocytes in vivo contributes to the understanding of the mechanism of the inflammatory process and to the development of anti-inflammatory drugs. Several roadblocks exist that hinder successful automated tracking including the moving background, the severe image noise and clutter, the occlusion of the target leukocyte by other leukocytes and structures, the jitter caused by the breathing movement of the living animal, and the weak image contrast. In this paper, a Monte Carlo tracker is developed for automatically tracking a single rolling leukocyte in vivo. Based on the leukocyte movement information and the image intensity features, a specialized sample-weighting criterion is tailored to the application. In comparison with a snake-based tracker, our experiments show that, as the noise intensity level increases, the performance of the snake tracker degrades more than that of the Monte Carlo tracker. In cases, where the leukocyte is observed in contact with the vessel wall, the Monte Carlo tracker is less affected by the image clutter. From tracking within 99 intravital microscopic video sequences, the Monte Carlo tracker exhibits superior performance in the reduced localization error and the increased number of frames tracked when compared with the centroid tracker, the correlation tracker and the GVF snake tracker.

Algorithms↗

Centerline-based colon segmentation for CT colonography.

We have developed a fully automated algorithm for colon segmentation, centerline-based segmentation (CBS), which is faster than any of the previously presented segmentation algorithms, but also has high sensitivity as well as high specificity. The algorithm first thresholds a set of unprocessed CT slices. Outer air is removed, after which a bounding box is computed. A centerline is computed for all remaining regions in the thresholded volume, disregarding segments related to extracolonic structures. Centerline segments are connected, after which the anatomy-based removal of segments representing extracolonic structures occurs. Segments related to the remaining centerline are locally region grown, and the colonic wall is found by dilation. Shape-based interpolation provides an isotropic mask. For 38 CT datasets, CBS was compared with the knowledge-guided segmentation (KGS) algorithm for sensitivity and specificity. With use of a 1.5 GHz AMD Athlon-based PC, the average computation time for the segmentation was 14.8 s. The sensitivity was, on average, 96%, and the specificity was 99%. A total of 21% of the voxels segmented by KGS, of which 96% represented extracolonic structures and 4% represented the colon, were removed.

Algorithms↗

Learning synergy in a multilevel neuronal architecture.

An artificial worlds model of the brain has been developed that integrates memory, intraneuronal dynamics and multilevel evolutionary learning. The model includes two major subsystems. The first is a memory-manipulation scheme, called the reference neuron system, that serves to orchestrate a repertoire of neurons with different input-output capabilities. Signals impinging on these neurons are integrated by a cytoskeletal structure that is simulated as a cellular automation. The second subsystem is an evolutionary learning scheme, called the selection circuits system, that serves to train the neurons in the repertoire by varying the cytoskeletal proteins that control signal flow or readouts. The integrated system comprises two layers of cytoskeletally controlled neurons and two layers of reference neurons. Evolution can occur at the level of readout enzymes in neurons, at the level of proteins that control the flow of signals in the cytoskeleton and at the level of reference neurons that orchestrate the repertoire. The integrated system controls the motion of a modeled organism that is embedded in an artificial environment consisting of barriers, food and a target. The organism effectively learns to use patterns of barriers in its local environment to find the target, using food as a reward. Experiments with the model show that the integrated system enjoys significant computational synergies that make it more powerful than the component systems standing alone, that interactions between different levels at which variation can occur exert significant control over the tempo of evolution, that the synergies between different components and levels becomes more important as the environment becomes more complex and that mutation strategies that significantly slow down the rate of learning significantly decrease the degrading effects of environmental noise on performance.

Biological Evolution↗

Ridge-based vessel segmentation in color images of the retina.

A method is presented for automated segmentation of vessels in two-dimensional color images of the retina. This method can be used in computer analyses of retinal images, e.g., in automated screening for diabetic retinopathy. The system is based on extraction of image ridges, which coincide approximately with vessel centerlines. The ridges are used to compose primitives in the form of line elements. With the line elements an image is partitioned into patches by assigning each image pixel to the closest line element. Every line element constitutes a local coordinate frame for its corresponding patch. For every pixel, feature vectors are computed that make use of properties of the patches and the line elements. The feature vectors are classified using a kappaNN-classifier and sequential forward feature selection. The algorithm was tested on a database consisting of 40 manually labeled images. The method achieves an area under the receiver operating characteristic curve of 0.952. The method is compared with two recently published rule-based methods of Hoover et al. and Jiang et al. The results show that our method is significantly better than the two rule-based methods (p < 0.01). The accuracy of our method is 0.944 versus 0.947 for a second observer.

Algorithms↗

The quest to deduce protein function from sequence: the role of pattern databases.

In the wake of the numerous now-fruitful genome projects, we have witnessed a 'tsunami' of sequence data and with it the birth of the field of bioinformatics. Bioinformatics involves the application of information technology to the management and analysis of biological data. For many of us, this means that databases and their search tools have become an essential part of the research environment. However, the rate of sequence generation and the haphazard proliferation of databases have made it difficult to keep pace with developments, even for the cognoscenti. Moreover, increasing amounts of sequence information do not necessarily equate with an increase in knowledge, and in the panic to automate the route from raw data to biological insight, we may be generating and propagating innumerable errors in our precious databases. In the genome era upon us, researchers want rapid, easy-to-use, reliable tools for functional characterisation of newly determined sequences. For the pharmaceutical industry in particular, the Pandora's box of bioinformatics harbours an information-rich nugget, ripe with potential drug targets and possible new avenues for the development of therapeutic agents. This review outlines the current status of the major pattern databases now used routinely in the analysis of protein sequences. The review is divided into three main sections. In the first, commonly used terms are defined and the methods behind the databases are briefly described; in the second, the structure and content of the principal pattern databases are discussed; and in the final part, several alignment databases, which are frequently confused with pattern databases, are mentioned. For the new-comer, the array of resources, the range of methods behind them and the different tools required to search them can be confusing. The review therefore also briefly mentions a current international endeavour to integrate the diverse databases, which effort should facilitate sequence analysis in the future. This is particularly important for target-discovery programmes, where the challenge is to rationalise the enormous numbers of potential targets generated by sequence database searches. This problem may be addressed, at least in part, by reducing search outputs to the more focused and manageable subsets suggested by searches of integrated groups of family-specific pattern databases.

Amino Acid Motifs↗

Knowledge discovery by automated identification and ranking of implicit relationships.

MOTIVATION: New relationships are often implicit from existing information, but the amount and growth of published literature limits the scope of analysis an individual can accomplish. Our goal was to develop and test a computational method to identify relationships within scientific reports, such that large sets of relationships between unrelated items could be sought out and statistically ranked for their potential relevance as a set. RESULTS: We first construct a network of tentative relationships between 'objects' of biomedical research interest (e.g. genes, diseases, phenotypes, chemicals) by identifying their co-occurrences within all electronically available MEDLINE records. Relationships shared by two unrelated objects are then ranked against a random network model to estimate the statistical significance of any given grouping. When compared against known relationships, we find that this ranking correlates with both the probability and frequency of object co-occurrence, demonstrating the method is well suited to discover novel relationships based upon existing shared relationships. To test this, we identified compounds whose shared relationships predicted they might affect the development and/or progression of cardiac hypertrophy. When laboratory tests were performed in a rodent model, chlorpromazine was found to reduce the progression of cardiac hypertrophy.

Abstracting and Indexing↗

A polygon matching algorithm and its applications to verification of radiation field placement in radiotherapy.

An automated method for quantitative evaluation of geometric accuracy of radiation field placement during radiation therapy treatments has been developed and is presented in this article. The method is based on matching polygons representing the boundaries of the prescribed and treated radiation fields. An efficient algorithm for finding the intersection of two arbitrary simple polygons has been developed for that purpose. The process of polygon matching is performed by iterative optimization of the goodness-of-match measure derived from the area of the intersection between the two polygons to be matched. The optimization is done with respect to the translation and rotation in the plane that are applied to the polygon representing the treated field. The optimum translation and rotation define the adjustments in the position of the treated field, relative to the patients anatomy, which are necessary in order to bring the treated field into the best possible agreement with the prescription. The radiation field matching scheme was applied to a clinical dataset consisting of prescription-treatment field pairs extracted from simulation and portal images for 40 patients. It is concluded that the observed accuracy and efficiency of the method make it very suitable for applications in radiation therapy verification.

Algorithms↗

Accurate detection of aneuploidies in array CGH and gene expression microarray data.

MOTIVATION: Chromosomal copy number changes (aneuploidies) are common in cell populations that undergo multiple cell divisions including yeast strains, cell lines and tumor cells. Identification of aneuploidies is critical in evolutionary studies, where changes in copy number serve an adaptive purpose, as well as in cancer studies, where amplifications and deletions of chromosomal regions have been identified as a major pathogenetic mechanism. Aneuploidies can be studied on whole-genome level using array CGH (a microarray-based method that measures the DNA content), but their presence also affects gene expression. In gene expression microarray analysis, identification of copy number changes is especially important in preventing aberrant biological conclusions based on spurious gene expression correlation or masked phenotypes that arise due to aneuploidies. Previously suggested approaches for aneuploidy detection from microarray data mostly focus on array CGH, address only whole-chromosome or whole-arm copy number changes, and rely on thresholds or other heuristics, making them unsuitable for fully automated general application to gene expression datasets. There is a need for a general and robust method for identification of aneuploidies of any size from both array CGH and gene expression microarray data. RESULTS: We present ChARM (Chromosomal Aberration Region Miner), a robust and accurate expectation-maximization based method for identification of segmental aneuploidies (partial chromosome changes) from gene expression and array CGH microarray data. Systematic evaluation of the algorithm on synthetic and biological data shows that the method is robust to noise, aneuploidal segment size and P-value cutoff. Using our approach, we identify known chromosomal changes and predict novel potential segmental aneuploidies in commonly used yeast deletion strains and in breast cancer. ChARM can be routinely used to identify aneuploidies in array CGH datasets and to screen gene expression data for aneuploidies or array biases. Our methodology is sensitive enough to detect statistically significant and biologically relevant aneuploidies even when expression or DNA content changes are subtle as in mixed populations of cells. AVAILABILITY: Code available by request from the authors and on Web supplement at http://function.cs.princeton.edu/ChARM/

Algorithms↗

Improved prediction of protein-protein binding sites using a support vector machines approach.

MOTIVATION: Structural genomics projects are beginning to produce protein structures with unknown function, therefore, accurate, automated predictors of protein function are required if all these structures are to be properly annotated in reasonable time. Identifying the interface between two interacting proteins provides important clues to the function of a protein and can reduce the search space required by docking algorithms to predict the structures of complexes. RESULTS: We have combined a support vector machine (SVM) approach with surface patch analysis to predict protein-protein binding sites. Using a leave-one-out cross-validation procedure, we were able to successfully predict the location of the binding site on 76% of our dataset made up of proteins with both transient and obligate interfaces. With heterogeneous cross-validation, where we trained the SVM on transient complexes to predict on obligate complexes (and vice versa), we still achieved comparable success rates to the leave-one-out cross-validation suggesting that sufficient properties are shared between transient and obligate interfaces. AVAILABILITY: A web application based on the method can be found at http://www.bioinformatics.leeds.ac.uk/ppi_pred. The dataset of 180 proteins used in this study is also available via the same web site. CONTACT: westhead@bmb.leeds.ac.uk SUPPLEMENTARY INFORMATION: http://www.bioinformatics.leeds.ac.uk/ppi-pred/supp-material.

Algorithms↗

An association rule mining-based methodology for automated detection of ischemic ECG beats.

Currently, an automated methodology based on association rules is presented for the detection of ischemic beats in long duration electrocardiographic (ECG) recordings. The proposed approach consists of three stages. 1) Preprocessing: Noise is removed and all the necessary ECG features are extracted. 2) Discretization: The continuous valued features are transformed to categorical. 3) CLASSIFICATION: An association rule extraction algorithm is utilized and a rule-based classification model is created. According to the proposed methodology, electrocardiogram (ECG) features extracted from the ST segment and the T-wave, as well as the patient's age, were used as inputs. The output was the classification of the beat as ischemic or not. Various algorithms were tested both for discretization and for classification using association rules. To evaluate the methodology, a cardiac beat dataset was constructed using several recordings of the European Society of Cardiology ST-T database. The obtained sensitivity (Se) and specificity (Sp) was 87% and 93%, respectively. The proposed methodology combines high accuracy with the ability to provide interpretation for the decisions made, since it is based on a set of association rules.

Arrhythmias, Cardiac↗

Evaluation of features for catalytic residue prediction in novel folds.

Structural genomics projects are determining the three-dimensional structure of proteins without full characterization of their function. A critical part of the annotation process involves appropriate knowledge representation and prediction of functionally important residue environments. We have developed a method to extract features from sequence, sequence alignments, three-dimensional structure, and structural environment conservation, and used support vector machines to annotate homologous and nonhomologous residue positions based on a specific training set of residue functions. In order to evaluate this pipeline for automated protein annotation, we applied it to the challenging problem of prediction of catalytic residues in enzymes. We also ranked the features based on their ability to discriminate catalytic from noncatalytic residues. When applying our method to a well-annotated set of protein structures, we found that top-ranked features were a measure of sequence conservation, a measure of structural conservation, a degree of uniqueness of a residue's structural environment, solvent accessibility, and residue hydrophobicity. We also found that features based on structural conservation were complementary to those based on sequence conservation and that they were capable of increasing predictor performance. Using a family nonredundant version of the ASTRAL 40 v1.65 data set, we estimated that the true catalytic residues were correctly predicted in 57.0% of the cases, with a precision of 18.5%. When testing on proteins containing novel folds not used in training, the best features were highly correlated with the training on families, thus validating the approach to nonhomologous catalytic residue prediction in general. We then applied the method to 2781 coordinate files from the structural genomics target pipeline and identified both highly ranked and highly clustered groups of predicted catalytic residues.

Algorithms↗

Registration of central paths and colonic polyps between supine and prone scans in computed tomography colonography: pilot study.

Computed tomography colonography (CTC) is a minimally invasive method that allows the evaluation of the colon wall from CT sections of the abdomen/pelvis. The primary goal of CTC is to detect colonic polyps, precursors to colorectal cancer. Because imperfect cleansing and distension can cause portions of the colon wall to be collapsed, covered with water, and/or covered with retained stool, patients are scanned in both prone and supine positions. We believe that both reading efficiency and computer aided detection (CAD) of CTC images can be improved by accurate registration of data from the supine and prone positions. We developed a two-stage approach that first registers the colonic central paths using a heuristic and automated algorithm and then matches polyps or polyp candidates (CAD hits) by a statistical approach. We evaluated the registration algorithm on 24 patient cases. After path registration, the mean misalignment distance between prone and supine identical anatomic landmarks was reduced from 47.08 to 12.66 mm, a 73% improvement. The polyp registration algorithm was specifically evaluated using eight patient cases for which radiologists identified polyps separately for both supine and prone data sets, and then manually registered corresponding pairs. The algorithm correctly matched 78% of these pairs without user input. The algorithm was also applied to the 30 highest-scoring CAD hits in the prone and supine scans and showed a success rate of 50% in automatically registering corresponding polyp pairs. Finally, we computed the average number of CAD hits that need to be manually compared in order to find the correct matches among the top 30 CAD hits. With polyp registration, the average number of comparisons was 1.78 per polyp, as opposed to 4.28 comparisons without polyp registration.

Algorithms↗

Automated functional classification of experimental and predicted protein structures.

BACKGROUND: Proteins that are similar in sequence or structure may perform different functions in nature. In such cases, function cannot be inferred from sequence or structural similarity. RESULTS: We analyzed experimental structures belonging to the Structural Classification of Proteins (SCOP) database and showed that about half of them belong to multi-functional fold families for which protein similarity alone is not adequate to assign function. We also analyzed predicted structures from the LiveBench and the PDB-CAFASP experiments and showed that accurate homology-based functional assignments cannot be achieved approximately one third of the time, when the protein is a member of a multi-functional fold family. We then conducted extended performance evaluation and comparisons on both experimental and predicted structures using our Functional Signatures from Structural Alignments (FSSA) algorithm that we previously developed to handle the problem of classifying proteins belonging to multi-functional fold families. CONCLUSION: The results indicate that the FSSA algorithm has better accuracy when compared to homology-based approaches for functional classification of both experimental and predicted protein structures, in part due to its use of local, as opposed to global, information for classifying function. The FSSA algorithm has also been implemented as a webserver and is available at http://protinfo.compbio.washington.edu/fssa.

Algorithms↗

Automated decomposition of intramuscular electromyographic signals.

We present a novel method for extracting and classifying motor unit action potentials (MUAPs) from one-channel electromyographic recordings. The extraction of MUAP templates is carried out using a symbolic representation of waveforms, a common technique in signature verification applications. The assignment of MUAPs to their specific trains is achieved by means of repeated template matching passes using pseudocorrelation, a new matched-filter-based similarity measure. Identified MUAPs are peeled off and the residual signal is analyzed using shortened templates to facilitate the resolution of superimpositions. The program was tested with simulated data and with experimental signals obtained using fine-wire electrodes in the biceps brachii during isometric contractions ranging from 5% to 30% of the maximum voluntary contraction. Analyzed signals were made of up to 14 MUAP trains. Most templates were extracted automatically, but complex signals sometimes required the adjustment of 2 parameters to account for all the MUAP trains present. Classification accuracy rates for simulations ranged from an average of 96.3% +/- 0.9% (4 trains) to 75.6% +/- 11.0% (12 trains). The classification portion of the program never required user intervention. Decomposition of most 10-s-long signals required less than 10 s using a conventional desktop computer, thus showing capabilities for real-time applications.

Action Potentials↗

The bi-elliptical deformable contour and its application to automated tongue segmentation in Chinese medicine.

Automated tongue image segmentation, in Chinese medicine, is difficult due to two special factors: 1) there are many pathological details on the surface of the tongue, which have a large influence on edge extraction; 2) the shapes of the tongue bodies captured from various persons (with different diseases) are quite different, so they are impossible to describe properly using a predefined deformable template. To address these problems, in this paper, we propose an original technique that is based on a combination of a bi-elliptical deformable template (BEDT) and an active contour model, namely the bi-elliptical deformable contour (BEDC). The BEDT captures gross shape features by using the steepest decent method on its energy function in the parameter space. The BEDC is derived from the BEDT by substituting template forces for classical internal forces, and can deform to fit local details. Our algorithm features fully automatic interpretation of tongue images and a consistent combination of global and local controls via the template force. We apply the BEDC to a large set of clinical tongue images and present experimental results.

Algorithms↗