Search PubMed⌕ Search

Biomedical subjects

Francisco Melo

Publications and source records attributed to Francisco Melo.

18 recordsLinked to original sources

Accurate and unambiguous tag-to-gene mapping in serial analysis of gene expression.

BACKGROUND: In this study, we present a robust and reliable computational method for tag-to-gene assignment in serial analysis of gene expression (SAGE). The method relies on current genome information and annotation, incorporation of several new features, and key improvements over alternative methods, all of which are important to determine gene expression levels more accurately. The method provides a complete annotation of potential virtual SAGE tags within a genome, along with an estimation of their confidence for experimental observation that ranks tags that present multiple matches in the genome. RESULTS: We applied this method to the Saccharomyces cerevisiae genome, producing the most thorough and accurate annotation of potential virtual SAGE tags that is available today for this organism. The usefulness of this method is exemplified by the significant reduction of ambiguous cases in existing experimental SAGE data. In addition, we report new insights from the analysis of existing SAGE data. First, we found that experimental SAGE tags mapping onto introns, intron-exon boundaries, and non-coding RNA elements are observed in all available SAGE data. Second, a significant fraction of experimental SAGE tags was found to map onto genomic regions currently annotated as intergenic. Third, a significant number of existing experimental SAGE tags for yeast has been derived from truncated cDNAs, which are synthesized through oligo-d(T) priming to internal poly-(A) regions during reverse transcription. CONCLUSION: We conclude that an accurate and unambiguous tag mapping process is essential to increase the quality and the amount of information that can be extracted from SAGE experiments. This is supported by the results obtained here and also by the large impact that the erroneous interpretation of these data could have on downstream applications.

Chromosome Mapping↗

A composite score for predicting errors in protein structure models.

Reliable prediction of model accuracy is an important unsolved problem in protein structure modeling. To address this problem, we studied 24 individual assessment scores, including physics-based energy functions, statistical potentials, and machine learning-based scoring functions. Individual scores were also used to construct approximately 85,000 composite scoring functions using support vector machine (SVM) regression. The scores were tested for their abilities to identify the most native-like models from a set of 6000 comparative models of 20 representative protein structures. Each of the 20 targets was modeled using a template of <30% sequence identity, corresponding to challenging comparative modeling cases. The best SVM score outperformed all individual scores by decreasing the average RMSD difference between the model identified as the best of the set and the model with the lowest RMSD (DeltaRMSD) from 0.63 A to 0.45 A, while having a higher Pearson correlation coefficient to RMSD (r=0.87) than any other tested score. The most accurate score is based on a combination of the DOPE non-hydrogen atom statistical potential; surface, contact, and combined statistical potentials from MODPIPE; and two PSIPRED/DSSP scores. It was implemented in the SVMod program, which can now be applied to select the final model in various modeling problems, including fold assignment, target-template alignment, and loop modeling.

Models, Molecular↗

Accuracy of sequence alignment and fold assessment using reduced amino acid alphabets.

Reduced or simplified amino acid alphabets group the 20 naturally occurring amino acids into a smaller number of representative protein residues. To date, several reduced amino acid alphabets have been proposed, which have been derived and optimized by a variety of methods. The resulting reduced amino acid alphabets have been applied to pattern recognition, generation of consensus sequences from multiple alignments, protein folding, and protein structure prediction. In this work, amino acid substitution matrices and statistical potentials were derived based on several reduced amino acid alphabets and their performance assessed in a large benchmark for the tasks of sequence alignment and fold assessment of protein structure models, using as a reference frame the standard alphabet of 20 amino acids. The results showed that a large reduction in the total number of residue types does not necessarily translate into a significant loss of discriminative power for sequence alignment and fold assessment. Therefore, some definitions of a few residue types are able to encode most of the relevant sequence/structure information that is present in the 20 standard amino acids. Based on these results, we suggest that the use of reduced amino acid alphabets may allow to increasing the accuracy of current substitution matrices and statistical potentials for the prediction of protein structure of remote homologs.

Amino Acid Sequence↗

Experimental evidence of shock mitigation in a Hertzian tapered chain.

We present an experimental study of the mechanical impulse propagation through a horizontal alignment of elastic spheres of progressively decreasing diameter phi(n): namely, a tapered chain. Experimentally, the diameters of spheres which interact via the Hertz potential are selected to keep as close as possible to an exponential decrease, phi(n+1) = (1-q)phi(n), where the experimental tapering factor is either q(1) approximately equal to 5.60% or q(2) approximately equal to 8.27%. In agreement with recent numerical results, an impulse initiated in a monodisperse chain (a chain of identical beads) propagates without shape changes and progressively transfers its energy and momentum to a propagating tail when it further travels in a tapered chain. As a result, the front pulse of this wave decreases in amplitude and accelerates. Both effects are satisfactorily described by the hard-sphere approximation, and basically, the shock mitigation is due to partial transmissions, from one bead to the next, of momentum and energy of the front pulse. In addition when small dissipation is included, better agreement with experiments is found. A close analysis of the loading part of the experimental pulses demonstrates that the front wave adopts a self-similar solution as it propagates in the tapered chain. Finally, our results corroborate the capability of these chains to thermalize propagating impulses and thereby act as shock absorbing devices.

Journal Article↗

MODBASE: a database of annotated comparative protein structure models and associated resources.

MODBASE (http://salilab.org/modbase) is a database of annotated comparative protein structure models for all available protein sequences that can be matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on MODELLER for fold assignment, sequence-structure alignment, model building and model assessment (http:/salilab.org/modeller). MODBASE is updated regularly to reflect the growth in protein sequence and structure databases, and improvements in the software for calculating the models. MODBASE currently contains 3 094 524 reliable models for domains in 1 094 750 out of 1 817 889 unique protein sequences in the UniProt database (July 5, 2005); only models based on statistically significant alignments and models assessed to have the correct fold despite insignificant alignments are included. MODBASE also allows users to generate comparative models for proteins of interest with the automated modeling server MODWEB (http://salilab.org/modweb). Our other resources integrated with MODBASE include comprehensive databases of multiple protein structure alignments (DBAli, http://salilab.org/dbali), structurally defined ligand binding sites and structurally defined binary domain interfaces (PIBASE, http://salilab.org/pibase) as well as predictions of ligand binding sites, interactions between yeast proteins, and functional consequences of human nsSNPs (LS-SNP, http://salilab.org/LS-SNP).

Binding Sites↗

Molecular and population analyses of a recombination event in the catabolic plasmid pJP4.

Cupriavidus necator JMP134(pJP4) harbors a catabolic plasmid, pJP4, which confers the ability to grow on chloroaromatic compounds. Repeated growth on 3-chlorobenzoate (3-CB) results in selection of a recombinant strain, which degrades 3-CB better but no longer grows on 2,4-dichlorophenoxyacetate (2,4-D). We have previously proposed that this phenotype is due to a double homologous recombination event between inverted repeats of the multicopies of this plasmid within the cell. One recombinant form of this plasmid (pJP4-F3) explains this phenotype, since it harbors two copies of the chlorocatechol degradation tfd gene clusters, which are essential to grow on 3-CB, but has lost the tfdA gene, encoding the first step in degradation of 2,4-D. The other recombinant plasmid (pJP4-FM) should harbor two copies of the tfdA gene but no copies of the tfd gene clusters. A molecular analysis using a multiplex PCR approach to distinguish the wild-type plasmid pJP4 from its two recombinant forms, was carried out. Expected PCR products confirming this recombination model were found and sequenced. Few recombinant plasmid forms in cultures grown in several carbon sources were detected. Kinetic studies indicated that cells containing the recombinant plasmid pJP4-FM were not selectable by sole carbon source growth pressure, whereas those cells harboring recombinant plasmid pJP4-F3 were selected upon growth on 3-CB. After 12 days of repeated growth on 3-CB, the complete plasmid population in C. necator JMP134 apparently corresponds to this form. However, wild-type plasmid forms could be recovered after growing this culture on 2,4-D, indicating that different plasmid forms can be found in C. necator JMP134 at the population level.

2,4-Dichlorophenoxyacetic Acid↗

Size segregation, convection, and arching effect.

We present a numerical study based on the contact dynamic procedure for the size segregation in a two-dimensional vibrating container. In agreement with previous studies, we find that the rising of the larger particle is accompanied by convection rolls. However, at low enough acceleration, rolls do not penetrate the entire cell and a local vault mechanism is observed, in which the larger particle moves up by steps. This mechanism, described in experiments by Duran and co-workers [Phys. Rev. E. 50, 5138 (1994)], occurs in a very narrow region of parameter space. Once the larger intruder reaches the exponential tail of convecting roll, its vertical motion becomes continuous and the corresponding rising speed is dramatically increased.

Journal Article↗

dnaMATE: a consensus melting temperature prediction server for short DNA sequences.

An accurate and robust large-scale melting temperature prediction server for short DNA sequences is dispatched. The server calculates a consensus melting temperature value using the nearest-neighbor model based on three independent thermodynamic data tables. The consensus method gives an accurate prediction of melting temperature, as it has been recently demonstrated in a benchmark performed using all available experimental data for DNA sequences within the length range of 16-30 nt. This constitutes the first web server that has been implemented to perform a large-scale calculation of melting temperatures in real time (up to 5000 DNA sequences can be submitted in a single run). The expected accuracy of calculations carried out by this server in the range of 50-600 mM monovalent salt concentration is that 89% of the melting temperature predictions will have an error or deviation of <5 degrees C from experimental data. The server can be freely accessed at http://dna.bio.puc.cl/tm.html. The standalone executable versions of this software for LINUX, Macintosh and Windows platforms are also freely available at the same web site. Detailed further information supporting this server is available at the same web site referenced above.

DNA↗

How hertzian solitary waves interact with boundaries in a 1D granular medium.

We perform measurements, numerical simulations, and quantitative comparisons with available theory on solitary wave propagation in a linear chain of beads without static preconstraint. By designing a nonintrusive force sensor to measure the impulse as it propagates along the chain, we study the solitary wave reflection at a wall. We show that the main features of solitary wave reflection depend on wall mechanical properties. Since previous studies on solitary waves have been performed at walls without these considerations, our experiment provides a more reliable tool to characterize solitary wave propagation. We find, for the first time, precise quantitative agreements.

Journal Article↗

Quantitative analysis of wine yeast gene expression profiles under winemaking conditions.

Wine fermentation is a dynamic and complex process in which the yeast cell is subjected to multiple stress conditions. A successful adaptation involves changes in gene expression profiles where a large number of genes are up- or downregulated. Functional genomic approaches are commonly used to obtain global gene expression profiles, thereby providing a comprehensive view of yeast physiology. We used SAGE to quantify gene expression profiles in an industrial strain of Saccharomyces cerevisiae under winemaking conditions. The transcriptome of wine yeast was analysed at three stages during the fermentation process, mid-exponential phase, and early- and late-stationary phases. Upon correlation with the yeast genome, we found three classes of transcripts: (a) sequences that corresponded to ORFs; (b) expressed sequences from intergenic regions; and (c) messengers that did not match the published reference yeast genome. In all fermentation phases studied, the most highly expressed genes related to energy production and stress response. For many pathways, including glycolysis, different transcript levels were observed during each phase. Different isoenzymes, including hexose transporters (HXT), were differentially induced, depending on the growth phase. About 10% of transcripts matched non-annotated ORF regions within the yeast genome and could correspond to small novel genes originally omitted in the first gene annotation effort. Up to 22% of transcripts, particularly at late-stationary phase, did not match any known location within the genome. As the available reference yeast genome was obtained from a laboratory strain, these expressed sequences could represent genes only expressed by an industrial yeast strain. Further studies are necessary to identify the role of these potential genes during wine fermentation.

Cluster Analysis↗

Adaptive evolution of the insulin gene in caviomorph rodents.

Insulin is a conservative molecule among mammals, maintaining both its structure and function. Rodents that belong to the Suborder Hystricognathi represent an exception, having a very divergent molecule with unusual physiological properties. In this work, we analyzed the evolutionary pattern of the insulin gene in caviomorph rodents (South American hystricomorph rodents). We found that these rodents have higher rates of nonsynonymous:synonymous substitutions (d(N)/d(S)) than nonhystricomorph rodents and that values are heterogeneous inside the group. We estimated codons under positive selection, specifically the second binding site (A13 and B17) and others related with hexamerization (B18, B20, and B22). In the monomer structure, all selected sites formed a single patch around the second binding site. In the hexamer structure, these amino acids were grouped into three major patches. In this structure, contacts between B chains involved all selected sites (except B18), and between faces in the center of the molecule, all contacts were among selected sites. While there is no clear hypothesis regarding the cause of this drastic change, experimental evidence does show that this group of rodents has some peculiarities in growth function, and, whether coincidental or not, these changes appeared together with important changes in life-history traits.

Adaptation, Physiological↗

Droplets of fine powders running uphill by vertical vibration.

We observe "droplets" forming when an inclined surface initially covered by fine powders is vibrated vertically. Droplets move uphill in the direction of maximum local slope. The speed of droplets is nearly independent on their size whereas it is an increasing function of the plate acceleration and inclination. By evacuating the container we show that the interstitial air flow plays an important role on droplets forming and their drift.

Computer Simulation↗

Comparison of different melting temperature calculation methods for short DNA sequences.

MOTIVATION: The overall performance of several molecular biology techniques involving DNA/DNA hybridization depends on the accurate prediction of the experimental value of a critical parameter: the melting temperature Tm. Till date, many computer software programs based on different methods and/or parameterizations are available for the theoretical estimation of the experimental Tm value of any given short oligonucleotide sequence. However, in most cases, large and significant differences in the estimations of Tm were obtained while using different methods. Thus, it is difficult to decide which Tm value is the accurate one. In addition, it seems that most people who use these methods are unaware about the limitations, which are well described in the literature but not stated properly or restricted the inputs of most of the web servers and standalone software programs that implement them. RESULTS: A quantitative comparison on the similarities and differences among some of the published DNA/DNA Tm calculation methods is reported. The comparison was carried out for a large set of short oligonucleotide sequences ranging from 16 to 30 nt long, which span the whole range of CG-content. The results showed that significant differences were observed in all the methods, which in some cases depend on the oligonucleotide length and CG-content in a non-trivial manner. Based on these results, the regions of consensus and disagreement for the methods in the oligonucleotide feature space were reported. Owing to the lack of sufficient experimental data, a fair and complete assessment of accuracy for the different methods is not yet possible. Inspite of this limitation, a consensus Tm with minimal error probability was calculated by averaging the values obtained from two or more methods that exhibit similar behavior to each particular combination of oligonucleotide length and CG-content class. Using a total of 348 DNA sequences in the size range between 16mer and 30mer, for which the experimental Tm data are available, we demonstrated that the consensus Tm is a robust and accurate measure. It is expected that the results of this work would be constituted as a useful set of guidelines to be followed for the successful experimental implementation of various molecular biology techniques, such as quantitative PCR, multiplex PCR and the design of optimal DNA microarrays.

Algorithms↗

Dynamics of developable cones under shear.

We identify and study a persistent structure characteristic of the post-buckling regime of a thin cylindrical shell subjected to axial torsion. It consists of a pair of developable cones ( d cones) joined by an S-shaped ridge, having a size of the order of the radius of the cylinder. We study its formation by applying a concentrated load at the center of the shell, which creates an isolated pair of d cones, joined by a straight ridge that progressively tilts when a torsion angle is imposed. We interpret this response as the equilibrium state of a pair of interacting d cones in the presence of an in-plane shear field, created by axial torsion, which tends to drive them away from each other. We find that the amplitude of displacement of the d cones for a given torsion angle is amplified by decreasing the thickness of the sheet, therefore concluding that the equilibrium state is the result of a balance between bending and stretching energies. We propose a model where the driving effect is the coupling between the deformation field around the d cones and the imposed shear field, while the stabilizing effect is the increasing bending energy of the system.

Journal Article↗

Experimental study of surface waves scattering by a single vortex and a vortex dipole.

Surface waves interacting with filamentary vortex offer an interesting tool to characterize static and dynamics of surface vorticity. An experimental study of the scattered wave by a single vortex as well as by a vortex dipole is reported. On a plane wave front, the vortex circulation introduces a spatial phase shift that gives rise to dislocated waves. Dislocations can be explained by the effect of the differential advection due to the vortex flow, on the propagating wave front. Both the Burgers vector of dislocations and the scattering cross section are measured in the deep water regime. The analogy between the wave-vortex interaction and the Aharonov-Bohm effect in quantum mechanics is explored by contrasting the Burgers vectors of dislocations as well as the form of the scattered wave in both cases. For the case of the hard core vortex, spiral waves are observed in agreement with theoretical works on both the Aharanov-Bohm effect and classical surface wave mechanics.

Journal Article↗

A tool to assist the study of specific features at protein binding sites.

The Protein Data Bank contains a large amount of proteins that have been solved with small ligands bound to them. This constitutes a rich source of information for the study of the specific requirements of protein sites to bind small molecules with a favorable free energy. The specific atomic composition and three-dimensional geometric restraints of protein binding sites for different ligands could be easily obtained from there. The development of accurate binding site descriptors in proteins constitutes a valuable tool to assist in the large-scale prediction and annotation of protein function in whole genomes. In this work, an integrated database containing some processed and calculated protein/ligand information is described. It is expected that this database will constitute a useful tool for people working in the prediction of protein function from its structure. The database is accessible from the Internet through a web server located at: http://protein.bio.puc.cl

Binding Sites↗

A novel extracellular multicopper oxidase from Phanerochaete chrysosporium with ferroxidase activity.

Lignin degradation by the white rot basidiomycete Phanerochaete chrysosporium involves various extracellular oxidative enzymes, including lignin peroxidase, manganese peroxidase, and a peroxide-generating enzyme, glyoxal oxidase. Recent studies have suggested that laccases also may be produced by this fungus, but these conclusions have been controversial. We identified four sequences related to laccases and ferroxidases (Fet3) in a search of the publicly available P. chrysosporium database. One gene, designated mco1, has a typical eukaryotic secretion signal and is transcribed in defined media and in colonized wood. Structural analysis and multiple alignments identified residues common to laccase and Fet3 sequences. A recombinant MCO1 (rMCO1) protein expressed in Aspergillus nidulans had a molecular mass of 78 kDa, as determined by sodium dodecyl sulfate-polyacrylamide gel electrophoresis, and the copper I-type center was confirmed by the UV-visible spectrum. rMCO1 oxidized various compounds, including 2,2'-azino(bis-3-ethylbenzthiazoline-6-sulfonate) (ABTS) and aromatic amines, although phenolic compounds were poor substrates. The best substrate was Fe2+, with a Km close to 2 micro M. Collectively, these results suggest that the P. chrysosporium genome does not encode a typical laccase but rather encodes a unique extracellular multicopper oxidase with strong ferroxidase activity.

Amino Acid Sequence↗

Statistical potentials for fold assessment.

A protein structure model generally needs to be evaluated to assess whether or not it has the correct fold. To improve fold assessment, four types of a residue-level statistical potential were optimized, including distance-dependent, contact, Phi/Psi dihedral angle, and accessible surface statistical potentials. Approximately 10,000 test models with the correct and incorrect folds were built by automated comparative modeling of protein sequences of known structure. The criterion used to discriminate between the correct and incorrect models was the Z-score of the model energy. The performance of a Z-score was determined as a function of many variables in the derivation and use of the corresponding statistical potential. The performance was measured by the fractions of the correctly and incorrectly assessed test models. The most discriminating combination of any one of the four tested potentials is the sum of the normalized distance-dependent and accessible surface potentials. The distance-dependent potential that is optimal for assessing models of all sizes uses both C(alpha) and C(beta) atoms as interaction centers, distinguishes between all 20 standard residue types, has the distance range of 30 A, and is derived and used by taking into account the sequence separation of the interacting atom pairs. The terms for the sequentially local interactions are significantly less informative than those for the sequentially nonlocal interactions. The accessible surface potential that is optimal for assessing models of all sizes uses C(beta) atoms as interaction centers and distinguishes between all 20 standard residue types. The performance of the tested statistical potentials is not likely to improve significantly with an increase in the number of known protein structures used in their derivation. The parameters of fold assessment whose optimal values vary significantly with model size include the size of the known protein structures used to derive the potential and the distance range of the accessible surface potential. Fold assessment by statistical potentials is most difficult for the very small models. This difficulty presents a challenge to fold assessment in large-scale comparative modeling, which produces many small and incomplete models. The results described in this study provide a basis for an optimal use of statistical potentials in fold assessment.

Algorithms↗