Search PubMed⌕ Search

Biomedical subjects

A Valencia

Publications and source records attributed to A Valencia.

At least 73 records · Page 4Linked to original sources

EUCLID: automatic classification of proteins in functional classes by their database annotations.

UNLABELLED: A tool is described for the automatic classification of sequences in functional classes using their database annotations. The Euclid system is based on a simple learning procedure from examples provided by human experts. AVAILABILITY: Euclid is freely available for academics at http://www.gredos.cnb.uam.es/EUCLID, with the corresponding dictionaries for the generation of three, eight and 14 functional classes. CONTACT: E-mail: valencia@cnb.uam.es SUPPLEMENTARY INFORMATION: The results of the EUCLID classification of different genomes are available at http://www.sander.ebi.ac. uk/genequiz/. A detailed description of the different applications mentioned in the text is available at http://www.gredos.cnb.uam. es/EUCLID/Full_Paper

Computational Biology↗

Automatic extraction of keywords from scientific text: application to the knowledge domain of protein families.

MOTIVATION: Annotation of the biological function of different protein sequences is a time-consuming process currently performed by human experts. Genome analysis tools encounter great difficulty in performing this task. Database curators, developers of genome analysis tools and biologists in general could benefit from access to tools able to suggest functional annotations and facilitate access to functional information. APPROACH: We present here the first prototype of a system for the automatic annotation of protein function. The system is triggered by collections of s related to a given protein, and it is able to extract biological information directly from scientific literature, i.e. MEDLINE abstracts. Relevant keywords are selected by their relative accumulation in comparison with a domain-specific background distribution. Simultaneously, the most representative sentences and MEDLINE abstracts are selected and presented to the end-user. Evolutionary information is considered as a predominant characteristic in the domain of protein function. Our system consequently extracts domain-specific information from the analysis of a set of protein families. RESULTS: The system has been tested with different protein families, of which three examples are discussed in detail here: 'ataxia-telangiectasia associated protein', 'ran GTPase' and 'carbonic anhydrase'. We found generally good correlation between the amount of information provided to the system and the quality of the annotations. Finally, the current limitations and future developments of the system are discussed. AVAILABILITY: The current system can be considered as a prototype system. As such, it can be accessed as a server at http://columba.ebi.ac. uk:8765/andrade/abx. The system accepts text related to the protein or proteins to be evaluated (optimally, the result of a MEDLINE search by keyword) and the results are returned in the form of Web pages for keywords, sentences and s. SUPPLEMENTARY INFORMATION: Web pages containing full information on the examples mentioned in the text are available at: http://www.cnb.uam.es/ approximately cnbprot/keywords/ CONTACT: valencia@cnb.uam.es

Abstracting and Indexing↗

Role of UEV-1, an inactive variant of the E2 ubiquitin-conjugating enzymes, in in vitro differentiation and cell cycle behavior of HT-29-M6 intestinal mucosecretory cells.

By means of differential RNA display, we have isolated a cDNA corresponding to transcripts that are down-regulated upon differentiation of the goblet cell-like HT-29-M6 human colon carcinoma cell line. These transcripts encode proteins originally identified as CROC-1 on the basis of their capacity to activate transcription of c-fos. We show that these proteins are similar in sequence, and in predicted secondary and tertiary structure, to the ubiquitin-conjugating enzymes, also known as E2. Despite the similarities, these proteins lack a critical cysteine residue essential for the catalytic activity of E2 enzymes and, in vitro, they do not conjugate or transfer ubiquitin to protein substrates. These proteins constitute a distinct subfamily within the E2 protein family and are highly conserved in phylogeny from yeasts to mammals. Therefore, we have designated them UEV (ubiquitin-conjugating E2 enzyme variant) proteins, defined as proteins similar in sequence and structure to the E2 ubiquitin-conjugating enzymes but lacking their enzymatic activity (HW/GDB-approved gene symbol, UBE2V). At least two human genes code for UEV proteins, and one of them, located on chromosome 20q13.2, is expressed as at least four isoforms, generated by alternative splicing. All human cell types analyzed expressed at least one of these isoforms. Constitutive expression of exogenous human UEV in HT-29-M6 cells inhibited their capacity to differentiate upon confluence and caused both the entry of a larger proportion of cells in the division cycle and an accumulation in G2-M. This was accompanied with a profound inhibition of the mitotic kinase, cdk1. These results suggest that UEV proteins are involved in the control of differentiation and could exert their effects by altering cell cycle distribution.

Amino Acid Sequence↗

Influenza virus epidemiological surveillance in Argentina, 1987-1993, with molecular characterization of 1990 and 1993 isolates.

This report describes findings from epidemiological surveillance of influenza virus in two cities in Argentina (Mar del Plata and Córdoba) from 1987 to 1993. It includes information on reporting and serologic characterization of isolated influenza viruses. In addition, determination was made of the nucleotide sequences of the HA1 subunits of five type A (subtype H3) viral strains isolated in the epidemics of 1990 and 1993. The incidence of illness, type of viruses isolated, and H gene sequences were similar to what has been reported from other parts of the world during the same period. The H3 strains isolated in the 1990 and 1993 seasons were somewhat removed in their molecular characteristics from the strains the World Health Organization recommended for vaccines for those years, and appeared closer to the strains recommended for vaccination in subsequent seasons.

Argentina↗

Are binding residues conserved?

We present our attempt to quantify the evolutionary dynamics of functional residues in a representative set of protein structures and their homologous sequences. Using the log-odds formalism, the preference for all twenty amino acids to be conserved or participate in binding (or active) sites is examined. It appears that while there is a tendency for functional residues to be conserved, the two preference scales do not coincide. Remarkable differences between amino acid types emerge from this comparative study. The current approach is expected to lead towards a better understanding of functional site architecture in proteins.

Amino Acid Sequence↗

Specific DNA recognition by the Aspergillus nidulans three zinc finger transcription factor PacC.

The three zinc fingers of PacC, the transcription factor mediating pH regulation in Aspergillus nidulans, are necessary and sufficient to recognise specifically the target ipnA2 site. Missing nucleoside footprints confirmed the core target (double-stranded) hexanucleotide 5'-GCCAAG-3'. Any base substitution resulted in substantial or complete loss of binding, excepting A5 (partially replaceable by G). A T preceding the hexanucleotide enhanced binding. Interference footprinting indicates that the four Gs and A4 participate in specific contacts and that five pyrimidines are essential for binding. The size of the target sequence and the amino acid sequence of finger 1 suggested that its probe helix would not participate in base-specific contacts. Using site-directed mutagenesis and analogy to GLI, we propose that finger 1 crucially interacts with finger 2, a pair of conserved Trp residues in the Cys knuckles contacting hydrophobically. Finger 2 would also participate in extensive base contacts with the 5' moiety of the hexanucleotide. The specificity mutation Lys159Gln shows that finger 3 binds the 3' moiety of the hexanucleotide. Replacement of residues in positions +3 (His128Asn) and +2 (Gln155Lys) of the reading helices of fingers 2 and 3, respectively, prevented binding. Our biochemical and molecular data plus modelling using previously determined zinc finger-DNA complexes, predict specific contacts of fingers 2 and 3 to ipnA2. Our data indicate compact organisation of the PacC-ipnA2 complex (with nearly every base involved in specific contacts), illustrate the binding versatility of zinc finger domains and should facilitate analysis of other PacC family members, including Saccharomyces cerevisiae RIM1.

Amino Acid Sequence↗

Correlated mutations contain information about protein-protein interaction.

Many proteins have evolved to form specific molecular complexes and the specificity of this interaction is essential for their function. The network of the necessary inter-residue contacts must consequently constrain the protein sequences to some extent. In other words, the sequence of an interacting protein must reflect the consequence of this process of adaptation. It is reasonable to assume that the sequence changes accumulated during the evolution of one of the interacting proteins must be compensated by changes in the other. Here we apply a method for detecting correlated changes in multiple sequence alignments to a set of interacting protein domains and show that positions where changes occur in a correlated fashion in the two interacting molecules tend to be close to the protein-protein interfaces. This leads to the possibility of developing a method for predicting contacting pairs of residues from the sequence alone. Such a method would not need the knowledge of the structure of the interacting proteins, and hence would be both radically different and more widely applicable than traditional docking methods. We indeed demonstrate here that the information about correlated sequence changes is sufficient to single out the right inter-domain docking solution amongst many wrong alternatives of two-domain proteins. The same approach is also used here in one case (haemoglobin) where we attempt to predict the interface of two different proteins rather than two protein domains. Finally, we report here a prediction about the inter-domain contact regions of the heat- shock protein Hsc70 based only on sequence information.

Amino Acid Sequence↗

DNA sequencing and analysis of 130 kb from yeast chromosome XV.

We have determined the nucleotide sequence of 129,524 bases of yeast (Saccharomyces cerevisiae) chromosome XV. Sequence analysis revealed the presence of 59 non-overlapping open reading frames (ORFs) of length > 300 bp, three tRNA genes, four delta elements and one Ty-element. Among the 21 previously known yeast genes (36% of all ORFs in this fragment) were nucleoporin (NUP1), ras protein (RAS1), RNA polymerase III (RPC1) and elongation factor 2 (EF2). Further, 31 ORFs (53% of the total) were found to be homologous to known protein or DNA sequences, or sequence patterns. For seven ORFs (11% of the total) no homology was found. Among the most interesting protein identification in this DNA fragment are an inositol polyphosphatase, the second gene of this type found in yeast (homologous to the human OCRL gene involved in Lowe's syndrome), a new ADP ribosylation factor of the arf6 subfamily, the first protein containing three C2 domains, and an ORF similar to a Bacillus subtilis cell-cycle related protein. For each ORF detailed sequence analysis was carried out, with a full consideration of its biological function and pointing out key regions of interest for further functional analysis.

ADP-Ribosylation Factors↗

A single residue substitution causes a switch from the dual DNA binding specificity of plant transcription factor MYB.Ph3 to the animal c-MYB specificity.

Transcription factor MYB.Ph3 from Petunia binds to two types of sequences, MBSI and MBSII, whereas murine c-MYB only binds to MBSI, and Am305 from Antirrhinum only binds to MBSII. DNA binding studies with hybrids of these proteins pointed to the N-terminal repeat (R2) as the most involved in determining binding to MBSI and/or MBSII, although some influence of the C-terminal repeat (R3) was also evident. Furthermore, a single residue substitution (Leu71 --> Glu) in MYB.Ph3 changed its specificity to that of c-MYB, and c-MYB with the reciprocal substitution (Glu132 --> Leu) essentially gained the MYB.Ph3 specificity. Molecular modeling and DNA binding studies with site-specific MYB.Ph3 mutants strongly supported the notion that the drastic changes in DNA binding specificity caused by the Leu --> Glu substitution reflect the fact that certain residues influence this property both directly, through base contacts, and indirectly, through interactions with other base-contacting residues, and that a single residue may establish alternative base contacts in different targets. Additionally, differential effects of mutations at non-base-contacting residues in MYB.Ph3 and c-MYB were observed, reflecting the importance of protein context on DNA binding properties of MYB proteins.

Amino Acid Sequence↗

Primitive interventricular septum, its primordium, and its contribution in the definitive interventricular septum: in vivo labelling study in the chick embryo heart.

BACKGROUND: Because the studies on the embryological development of the primitive interventricular septum have been done with postmortem material, we do not know the site within the cardiac tube and the developmental stage at which the primordium appears and its anatomical manifestation in the mature heart. Consequently, we do not know its real contribution to the constitution of the definitive interventricular septum. METHODS: With this purpose, we selected an adequate biological model, the chick embryo heart and the in vivo labelling technique. We placed a label of gelatin India ink in the ventral fusion line of both cardiac primordia at the level of the interventricular grooves in the straight tube heart (stage 9+HH), and we traced the ink up to the mature heart (stage 36HH). We made histological sections of some hearts, of the zone where the label was found to investigate the first morphological manifestation of the primitive interventricular septum. We also made microdissections and scanning electron microscopic studies. RESULTS: The label placed at stage 9+HH in the ventral fusion line of both cardiac primordia, at the level of the interventricular grooves, was found at stage 14HH in the greater curvature of the looped heart, opposite the left interventricular groove. This label at stage 17HH was found in the apical trabecular region of the first cardiac septum (8-shaped septum) and in the mature heart (stage 36HH) in the definitive interventricular septum at the limit between the basal and the medial third of the definitive interventricular septum. CONCLUSIONS: Firstly, the primordium of the primitive interventricular septum appears at stage 9+HH, in the ventral fusion line of both cardiac primordia at the level of the interventricular grooves. Secondly, its first morphological manifestation takes place at stage 17HH, and it forms the apical trabeculated region of the first cardiac septum (8-shaped septum). Thirdly, the primitive interventricular septum gives origin to the middle and apical third of the definitive interventricular septum.

Animals↗

Conserved clusters of functionally related genes in two bacterial genomes.

An approach for genome comparison, combining function classification of gene products and sequence comparison, is presented. The genomes of Haemophilus influenzae and Escherichia coli are analyzed, and all genes are classified into nine major functional classes, corresponding to important cellular processes. To study gene order relationships and genome organization in the two bacteria, we performed statistics on neighboring pairs of genes. To estimate the significance of the observations, a statistical model based on binomial distributions has been developed. Significant patterns of gene order are observed within, as well as between, the two bacterial genomes: Functionally related genes tend to be neighbors more often than do unrelated genes. Some of these groups represent well-known operons, but additional gene clusters are identified. These clusters correspond to genomic elements that have been conserved during bacterial evolution. In addition to nearest-neighbor relationships, the method is also useful to study the relative direction of transcription in genomes, which is also highly conserved between homologous gene pairs. This new approach combines the high-level description of molecular function with pair statistics that express genome organization. It is expected to complement traditional methods of sequence analysis in the study of genomic structure, function, and evolution.

Binomial Distribution↗

Classification of protein families and detection of the determinant residues with an improved self-organizing map.

Using a SOM (self-organizing map) we can classify sequences within a protein family into subgroups that generally correspond to biological subcategories. These maps tend to show sequence similarity as proximity in the map. Combining maps generated at different levels of resolution, the structure of relations in protein families can be captured that could not otherwise be represented in a single map. The underlying representation of maps enables us to retrieve characteristic sequence patterns for individual subgroups of sequences. Such patterns tend to correspond to functionally important regions. We present a modified SOM algorithm that includes a convergence test that dynamically controls the learning parameters to adapt them to the learning set instead of being fixed and externally optimized by trial and error. Given the variability of protein family size and distribution, the addition of this features is necessary. The method is successfully tested with a number of families. The rab family of small GTPases is used to illustrate the performance of the method.

Algorithms↗

A novel common missense mutation G301C in the N-acetylgalactosamine-6-sulfate sulfatase gene in mucopolysaccharidosis IVA.

Mucopolysaccharidosis IVA (MPS IVA) is an autosomal recessive lysosomal storage disorder caused by a genetic defect in N-acetylgalactosamine-6-sulfate sulfatase (GALNS). In previous studies, we have found two common mutations in Caucasians and Japanese, respectively. To characterize the mutational spectrum in various ethnic groups, mutations in the GALNS gene in Colombian MPS IVA patients were investigated, and genetic backgrounds were extensively analyzed to identify racial origin, based on mitochondrial DNA (mtDNA) lineages. Three novel missense mutations never identified previously in other populations and found in 16 out of 19 Colombian MPS IVA unrelated alleles account for 84.2% of the alleles in this study. The G301C and S162F mutations account for 68.4% and 10.5% of mutations, respectively, whereas the remaining F69V is limited to a single allele. The skewed prevalence of G301C in only Colombian patients and haplotype analysis by restriction fragment length polymorphisms in the GALNS gene suggest that G301C originated from a common ancestor. Investigation of the genetic background by means of mtDNA lineages indicate that all our patients are probably of native American descent.

Asian People↗

Improving contact predictions by the combination of correlated mutations and other sources of sequence information.

We have previously developed a method for predicting interresidue contacts using information about correlated mutations in multiple sequence alignments. The predictions generated with this method were clearly better than random but not enough for their use in de novo protein folding experiments. We assess the possibility of improving contact predictions combining information from the following variables: correlated mutations, sequence conservation, sequence separation along the chain, alignment stability, family size, residue-specific contact occupancy and formation of contact networks. The application of a protocol for combining these independent variables leads to contact predictions that are on average two times better than those obtained initially with correlated mutations. Correlated mutations can be effectively combined with other types of information derived from multiple sequence alignments. Among the different variables tried, sequence conservation and contact density are particularly relevant for the combination with correlated mutations.

Amino Acid Sequence↗

Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.

We have developed a prototype for the automatic annotation of functional characteristics in protein families. The system is able to extract biological information directly from scientific literature in the form of MEDLINE abstracts. The criterion for selecting relevant keywords is the difference between their frequency in the abstracts associated with the protein family under study and its frequency in other unrelated protein families. The concept of functional information associated to protein families is the key feature of our system and gathers evolutionary information into the problem of functional annotation of biological sequences. The system has been tested in two different scenarios: first, a large set of protein families with a small number of abstract per family and second, selected protein families with large number of abstracts attached to each one. In both cases the performances are compared with annotations provided by human experts showing a clear relation between the amount of information provided to the system and the quality of the annotations. The automatic annotations are in many cases of similar quality to the ones contained in current data bases. The possibilities and difficulties to be encountered during the development of a full system for automatic annotation are discussed.

Abstracting and Indexing↗