Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Pattern Recognition, Automated”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Domain discovery method for topological profile searches in protein structures.

We describe a method for automated domain discovery for topological profile searches in protein structures. The method is used in a system TOPStructure for fast prediction of CATH classification for protein structures (given as PDB files). It is important for profile searches in multi-domain proteins, for which the profile method by itself tends to perform poorly. We also present an O(C(n)k + nk(2)) time algorithm for this problem, compared to the O(C(n)k + (nk)(2)) time used by a trivial algorithm (where n is the length of the structure, k is the number of profiles and C(n) is the time needed to check for a presence of a given motif in a structure of length n). This method has been developed and is currently used for TOPS representations of protein structures and prediction of CATH classification, but may be applied to other graph-based representations of protein or RNA structures and/or other prediction problems. A protein structure prediction system incorporating the domain discovery method is available at http://bioinf.mii.lu.lv/tops/.

Algorithms↗

Automatic scoring and quality assessment using accuracy bounds for FP-TDI SNP genotyping data.

BACKGROUND: Human diversity, namely single nucleotide polymorphisms (SNPs), is becoming a focus of biomedical research. Despite the binary nature of SNP determination, the majority of genotyping assay data need a critical evaluation for genotype calling. We applied statistical models to improve the automated analysis of 2-dimensional SNP data. METHODS: We derived several quantities in the framework of Gaussian mixture models that provide figures of merit to objectively measure the data quality. The accuracy of individual observations is scored as the probability of belonging to a certain genotype cluster, while the assay quality is measured by the overlap between the genotype clusters. RESULTS: The approach was extensively tested with a dataset of 438 nonredundant SNP assays comprising >150,000 datapoints. The performance of our automatic scoring method was compared with manual assignments. The agreement for the overall assay quality is remarkably good, and individual observations were scored differently by man and machine in 2.6% of cases, when applying stringent probability threshold values. CONCLUSION: Our definition of bounds for the accuracy for complete assays in terms of misclassification probabilities goes beyond other proposed analysis methods. We expect the scoring method to minimise human intervention and provide a more objective error estimate in genotype calling.

Algorithms↗

Automated 2-D cephalometric analysis on X-ray images by a model-based approach.

Craniofacial landmark localization and anatomical structure tracing on cephalograms are two important ways to obtain the cephalometric analysis. In order to computerize them in parallel, a model-based approach is proposed to locate 262 craniofacial feature points, including 90 landmarks and 172 auxiliary points. In model training, 12 landmarks are selected as reference points and used to divide every training shape to 10 regions according to the anatomical knowledge; principle components analysis is employed to characterize the region shape variations and the statistical grey profile of every feature point. Locating feature points on an input image is a two-stage procedure. First, we identify the reference landmarks by image processing and pattern matching techniques, so that the shape partition is performed on the input image. Then, for each region, its feature points are located by a modified active shape model. All craniofacial anatomical structures can be traced out by connecting the located points with subdivision curves according to the prior knowledge. Users are permitted to modify the results interactively in many different ways. Experimental results show the advantage and reliability of the proposed method.

Adolescent↗

Telephony-based voice pathology assessment using automated speech analysis.

A system for remotely detecting vocal fold pathologies using telephone-quality speech is presented. The system uses a linear classifier, processing measurements of pitch perturbation, amplitude perturbation and harmonic-to-noise ratio derived from digitized speech recordings. Voice recordings from the Disordered Voice Database Model 4337 system were used to develop and validate the system. Results show that while a sustained phonation, recorded in a controlled environment, can be classified as normal or pathologic with accuracy of 89.1%, telephone-quality speech can be classified as normal or pathologic with an accuracy of 74.2%, using the same scheme. Amplitude perturbation features prove most robust for telephone-quality speech. The pathologic recordings were then subcategorized into four groups, comprising normal, neuromuscular pathologic, physical pathologic and mixed (neuromuscular with physical) pathologic. A separate classifier was developed for classifying the normal group from each pathologic subcategory. Results show that neuromuscular disorders could be detected remotely with an accuracy of 87%, physical abnormalities with an accuracy of 78% and mixed pathology voice with an accuracy of 61%. This study highlights the real possibility for remote detection and diagnosis of voice pathology.

Algorithms↗

A method for the automated detection of venous gas bubbles in humans using empirical mode decomposition.

Doppler ultrasound signals are widely used to grade the quantity of circulating venous bubbles in divers. Current techniques rely on trained observers, making the grading process both time-consuming and subjective. The automated detection of bubbles, however, is confounded by the presence of other signals, primarily those arising from blood motion. Empirical Mode Decomposition was used here to calculate the intrinsic mode functions (IMFs) of a number of Doppler ultrasound signals from recreational divers, post-decompression. The IMFs provide a basis set for signal decomposition, each IMF corresponding to a different timescale in the signal. Each signal was found to comprise approximately 20 IMFs: the precise number being dependent upon the nature of the signal. A method is presented to detect bubbles using the IMF; features are first identified in the individual heart cycles, these having been previously determined using a robust peak detection method, by examining deviations from the ensemble averaged IMF. Bubbles are then identified as features appearing in more than one IMF, with significant energy in the original signal. This method has been applied to a subset of the available database and appears to perform with good sensitivity even when the signal has variable signal strength.

Artificial Intelligence↗

RotoTexture: automated tools for texturing raw video.

We propose a video editing system that allows a user to apply a time-coherent texture to a surface depicted in the raw video from a single uncalibrated camera, including the surface texture mapping of a texture image and the surface texture synthesis from a texture swatch. Our system avoids the construction of a 3D shape model and instead uses the recovered normal field to deform the texture so that it plausibly adheres to the undulations of the depicted surface. The texture mapping method uses the nonlinear least-squares optimization of a spring model to control the behavior of the texture image as it is deformed to match the evolving normal field through the video. The texture synthesis method uses a coarse optical flow to advect clusters of pixels corresponding to patches of similarly oriented surface points. These clusters are organized into a minimum advection tree to account for the dynamic visibility of clusters. We take a rather crude approach to normal recovering and optical flow estimation, yet the results are robust and plausible for nearly diffuse surfaces such as faces and t-shirts.

Algorithms↗

Using the fast marching method to extract curves with given global properties.

Curves are often used as anatomical features to match surfaces that represent biological objects, such as the human brain. Automated and semi-automated methods for extracting these curves usually rely on local properties of the surfaces such as the mean surface curvature without considering the global appearance of the curves themselves. These methods may require additional human intervention, and sometimes produce erroneous results. In this paper, we present an algorithm that is based on the fast marching method (FMM) to extract weighted geodesic curves. Instead of directly using the local image properties as a weight function, we use the surface properties, together with the global properties of the curves, to compute a weight function. This weight function is then used by the FMM to extract curves between given points. The general framework can be used to extract curves with different global properties. The resulting curves are guaranteed to be weighted geodesic curves without cusps usually introduced by intermediate points through which the curves are forced to pass. We show some results on both a simulated image and a highly convoluted human brain cortical surface.

Algorithms↗

Automatic registration of microarray images. I. Rectangular grid.

MOTIVATION: The analysis of high-throughput experiment data provided by microarrays becomes increasingly more and more important part of modern biological science. Microarrays allow to conduct genotyping or gene expression experiments on hundreds of thousands of test genes in parallel. Because of the large and constantly growing amount of experimental data the necessity of efficiency, robustness and complete automation of microarray image analysis algorithms is gaining significant attention in the field of microarray processing. RESULTS: The author presents here an efficient and completely automatic image registration algorithm (that is an algorithm for spots and blocks indexing) that allows to process a wide variety of microarray slides with different parameters of grid and block spacing as well as spot sizes. The algorithm scales linearly with the grid size, the time complexity is O(M), where M is number of rows x number of columns. It can successfully cope with local and global distortions of the grid, such as focal distortions and non-orthogonal transformations. The algorithm has been tested both on CCD and scanned images and showed very good performance-the processing time of a single slide with 44 blocks of 200 x 200 grid points (or 1 760 000 grid points total) was about 10 s. AVAILABILITY: The test implementation of the algorithm will be available upon request for academics. SUPPLEMENTARY INFORMATION: http://fleece.ucsd.edu/~vit/Registration_Supplement.pdf

Algorithms↗

Correspondence-free determination of the affine fundamental matrix.

Fundamental matrix estimation is a central problem in computer vision and forms the basis of tasks such as stereo imaging and structure from motion. Existing algorithms typically analyze the relative geometries of matched feature points identified in both projected views. Automated feature matching is itself a challenging problem. Results typically have a large number of false matches. Traditional fundamental matrix estimation methods are very sensitive to matching errors, which led naturally to the application of robust statistical estimation techniques to the problem. In this work, an entirely novel approach is proposed to the fundamental matrix estimation problem. Instead of analyzing the geometry of matched feature points, the problem is recast in the frequency domain through the use of Integral Projection, showing how this is a reasonable model for orthographic cameras. The problem now reduces to one of identifying matching lines in the frequency domain which, most importantly, requires no feature matching or correspondence information. Experimental results on both real and synthetic data are presented that demonstrate the algorithm is a practical technique for fundamental matrix estimation. The behavior of the proposed algorithm is additionally characterized with respect to input noise, feature counts, and other parameters of interest.

Algorithms↗

Automated identification of single nucleotide polymorphisms from sequencing data.

The single nucleotide polymorphism (SNP) is the difference of the DNA sequence between individuals and provides abundant information about genetic variation. Large scale discovery of high frequency SNPs is being undertaken using various methods. However, the publicly available SNP data sometimes need to be verified. If only a particular gene locus is concerned, locus-specific polymerase chain reaction amplification may be useful. Problem of this method is that the secondary peak has to be measured. We have analyzed trace data from conventional sequencing equipment and found an applicable rule to discern SNPs from noise. The rule is applied to multiply aligned sequences with a trace and the peak height of the traces are compared between samples. We have developed software that integrates this function to automatically identify SNPs. The software works accurately for high quality sequences and also can detect SNPs in low quality sequences. Further, it can determine allele frequency, display this information as a bar graph and assign corresponding nucleotide combinations. It is also designed for a person to verify and edit sequences easily on the screen. It is very useful for identifying de novo SNPs in a DNA fragment of interest.

Algorithms↗

Objective detection of evoked potentials using a bootstrap technique.

Evoked potentials are usually evaluated subjectively, by visual inspection, and considerable differences between interpretations can occur. Objective, automated methods are normally based on calculating one (or more) parameters from the data, but only some of these techniques can provide statistical significance (p-values) for the presence of a response. In this work, we propose a bootstrap technique to provide such p-values, which can be applied to a wide variety of parameters. The bootstrap method is based on randomly resampling (with replacement) the original data and gives an estimate of the probability that the response obtained is due to random variation in the data rather than a physiological response. The method is illustrated using auditory brainstem responses (ABRs) to detecting hearing thresholds. The flexibility of the approach is illustrated, showing how it can be used with different parameters, numbers of stimuli and with user-defined false-positive rates. The bootstrap method provides a new, simple and yet powerful means of detecting evoked potentials, which is very flexible and readily adapted to a wide variety of signal parameters.

Adolescent↗

Automated protein classification using consensus decision.

We propose a novel technique for automatically generating the SCOP classification of a protein structure with high accuracy. High accuracy is achieved by combining the decisions of multiple methods using the consensus of a committee (or an ensemble) classifier. Our technique is rooted in machine learning which shows that by judicially employing component classifiers, an ensemble classifier can be constructed to outperform its components. We use two sequence- and three structure-comparison tools as component classifiers. Given a protein structure, using the joint hypothesis, we first determine if the protein belongs to an existing category (family, superfamily, fold) in the SCOP hierarchy. For the proteins that are predicted as members of the existing categories, we compute their family-, superfamily-, and fold-level classifications using the consensus classifier. We show that we can significantly improve the classification accuracy compared to the individual component classifiers. In particular, we achieve error rates that are 3-12 times less than the individual classifiers' error rates at the family level, 1.5-4.5 times less at the superfamily level, and 1.1-2.4 times less at the fold level.

Algorithms↗

IsoSVM--distinguishing isoforms and paralogs on the protein level.

BACKGROUND: Recent progress in cDNA and EST sequencing is yielding a deluge of sequence data. Like database search results and proteome databases, this data gives rise to inferred protein sequences without ready access to the underlying genomic data. Analysis of this information (e.g. for EST clustering or phylogenetic reconstruction from proteome data) is hampered because it is not known if two protein sequences are isoforms (splice variants) or not (i.e. paralogs/orthologs). However, even without knowing the intron/exon structure, visual analysis of the pattern of similarity across the alignment of the two protein sequences is usually helpful since paralogs and orthologs feature substitutions with respect to each other, as opposed to isoforms, which do not. RESULTS: The IsoSVM tool introduces an automated approach to identifying isoforms on the protein level using a support vector machine (SVM) classifier. Based on three specific features used as input of the SVM classifier, it is possible to automatically identify isoforms with little effort and with an accuracy of more than 97%. We show that the SVM is superior to a radial basis function network and to a linear classifier. As an example application we use IsoSVM to estimate that a set of Xenopus laevis EST clusters consists of approximately 81% cases where sequences are each other's paralogs and 19% cases where sequences are each other's isoforms. The number of isoforms and paralogs in this allotetraploid species is of interest in the study of evolution. CONCLUSION: We developed an SVM classifier that can be used to distinguish isoforms from paralogs with high accuracy and without access to the genomic data. It can be used to analyze, for example, EST data and database search results. Our software is freely available on the Web, under the name IsoSVM.

Algorithms↗

Tracking of non-rigid articulatory organs in X-ray image sequences.

This article presents a system for the automated tracking of non-rigid anatomic structures in two-dimensional image sequences, which was primarily applied to X-ray image sequences of the vocal tract. In this particular application articulatory organs have to be measured to investigate the complex dynamic characteristics of human speech production. Of particular interest is a robust boundary detection of non-rigid organs such as lips and tongue. To solve this ill-posed detection problem under the presence of transparently superimposing structures, varying textural appearances of organs and noise, a two-level system is proposed. At the lower level, several edge-, region-, and motion-based image operators are combined to exploit their respective benefits and concomitantly compensate for their deficiencies. For the sake of precision, the result of these operators are not represented as larger tokens, such as line segments, but remain pixel-related cues or image evidences. At the higher level, an active contour-based component allows for the introduction of a priori knowledge about the object to be detected.

Algorithms↗

Bayesian search of functionally divergent protein subgroups and their function specific residues.

MOTIVATION: The rapid increase in the amount of protein sequence data has created a need for an automated identification of evolutionarily related subgroups from large datasets. The existing methods typically require a priori specification of the number of putative groups, which defines the resolution of the classification solution. RESULTS: We introduce a Bayesian model-based approach to simultaneous identification of evolutionary groups and conserved parts of the protein sequences. The model-based approach provides an intuitive and efficient way of determining the number of groups from the sequence data, in contrast to the ad hoc methods often exploited for similar purposes. Our model recognizes the areas in the sequences that are relevant for the clustering and regards other areas as noise. We have implemented the method using a fast stochastic optimization algorithm which yields a clustering associated with the estimated maximum posterior probability. The method has been shown to have high specificity and sensitivity in simulated and real clustering tasks. With real datasets the method also highlights the residues close to the active site. AVAILABILITY: Software 'kPax' is available at http://www.rni.helsinki.fi/jic/softa.html

Algorithms↗

Segmentation methodology for automated classification and differentiation of soft tissues in multiband images of high-resolution ultrasonic transmission tomography.

This paper presents a novel segmentation methodology for automated classification and differentiation of soft tissues using multiband data obtained with the newly developed system of high-resolution ultrasonic transmission tomography (HUTT) for imaging biological organs. This methodology extends and combines two existing approaches: the L-level set active contour (AC) segmentation approach and the agglomerative hierarchical kappa-means approach for unsupervised clustering (UC). To prevent the trapping of the current iterative minimization AC algorithm in a local minimum, we introduce a multiresolution approach that applies the level set functions at successively increasing resolutions of the image data. The resulting AC clusters are subsequently rearranged by the UC algorithm that seeks the optimal set of clusters yielding the minimum within-cluster distances in the feature space. The presented results from Monte Carlo simulations and experimental animal-tissue data demonstrate that the proposed methodology outperforms other existing methods without depending on heuristic parameters and provides a reliable means for soft tissue differentiation in HUTT images.

Algorithms↗

Quantitative assessment of liver fibrosis: a novel automated image analysis method.

BACKGROUND: Semiquantitative staging of liver fibrosis is a highly subjective procedure and may lead to an uncertainty in judgment regarding the degree of severity and hence the progression of the disease. AIM: In this work, we present an automated quantification system (AQS) for evaluating the degree of severity of fibrosis in liver biopsies based on Ishak et al.'s classification. Accordingly, liver fibrosis is classified into six classes depending on its severity and progression. The described system is of special value in accurately assessing the prognosis of chronic liver disease. METHODS: In our method, we tried to approximate the architecture of the fibrosis in the subject sample using texture features and shape representation of the fibrosis structural expansion with an overall accuracy of about 98%. RESULTS AND CONCLUSION: The presented AQS is considered to be a novel approach in the domain of automatic liver fibrosis quantification. It is a true quantification and intelligent approach that attempts to utilize the current semiquantitative methods of liver fibrosis assessment to turn them into real quantitative ones with significant reduction in variability and subjectivity. We propose that our method can be adopted by a panel of expert liver pathologists and software to be developed and used on a wide scale.

Biopsy↗

Feature selection and transduction for prediction of molecular bioactivity for drug design.

MOTIVATION: In drug discovery a key task is to identify characteristics that separate active (binding) compounds from inactive (non-binding) ones. An automated prediction system can help reduce resources necessary to carry out this task. RESULTS: Two methods for prediction of molecular bioactivity for drug design are introduced and shown to perform well in a data set previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001. The data is characterized by very few positive examples, a very large number of features (describing three-dimensional properties of the molecules) and rather different distributions between training and test data. Two techniques are introduced specifically to tackle these problems: a feature selection method for unbalanced data and a classifier which adapts to the distribution of the the unlabeled test data (a so-called transductive method). We show both techniques improve identification performance and in conjunction provide an improvement over using only one of the techniques. Our results suggest the importance of taking into account the characteristics in this data which may also be relevant in other problems of a similar type.

Algorithms↗