Search PubMed⌕ Search

Biomedical subjects

Ildefonso Cases

Publications and source records attributed to Ildefonso Cases.

13 recordsLinked to original sources

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software↗

CoGenT++: an extensive and extensible data environment for computational genomics.

MOTIVATION: CoGenT++ is a data environment for computational research in comparative and functional genomics, designed to address issues of consistency, reproducibility, scalability and accessibility. DESCRIPTION: CoGenT++ facilitates the re-distribution of all fully sequenced and published genomes, storing information about species, gene names and protein sequences. We describe our scalable implementation of ProXSim, a continually updated all-against-all similarity database, which stores pairwise relationships between all genome sequences. Based on these similarities, derived databases are generated for gene fusions--AllFuse, putative orthologs--OFAM, protein families--TRIBES, phylogenetic profiles--ProfUse and phylogenetic trees. Extensions based on the CoGenT++ environment include disease gene prediction, pattern discovery, automated domain detection, genome annotation and ancestral reconstruction. CONCLUSION: CoGenT++ provides a comprehensive environment for computational genomics, accessible primarily for large-scale analyses as well as manual browsing.

Chromosome Mapping↗

Promoters in the environment: transcriptional regulation in its natural context.

Transcriptional activation of many bacterial promoters in their natural environment is not a simple on/off decision. The expression of cognate genes is integrated in layers of iterative regulatory networks that ensure the performance not only of the whole cell, but also of the bacterial population, and even the microbial community, in a changing environment. Unlike in vitro systems, where transcription initiation can be recreated with a handful of essential components, in vivo, promoters must process various physicochemical and metabolic signals to determine their output. This helps to achieve optimal bacterial fitness in extremely competitive niches. Promoters therefore merge specific responses to distinct signals with inclusive reactions to more general environmental changes.

Adaptation, Physiological↗

BioLayout(Java): versatile network visualisation of structural and functional relationships.

Visualisation of biological networks is becoming a common task for the analysis of high-throughput data. These networks correspond to a wide variety of biological relationships, such as sequence similarity, metabolic pathways, gene regulatory cascades and protein interactions. We present a general approach for the representation and analysis of networks of variable type, size and complexity. The application is based on the original BioLayout program (C-language implementation of the Fruchterman-Rheingold layout algorithm), entirely re-written in Java to guarantee portability across platforms. BioLayout(Java) provides broader functionality, various analysis techniques, extensions for better visualisation and a new user interface. Examples of analysis of biological networks using BioLayout(Java) are presented.

Computer Graphics↗

Genetically modified organisms for the environment: stories of success and failure and what we have learned from them.

The expectations raised in the mid-1980s on the potential of genetic engineering for in situ remediation of environmental pollution have not been entirely fulfilled. Yet, we have learned a good deal about the expression of catabolic pathways by bacteria in their natural habitats, and how environmental conditions dictate the expression of desired catalytic activities. The many different choices between nutrients and responses to stresses form a network of transcriptional switches which, given the redundance and robustness of the regulatory circuits involved, can be neither unraveled through standard genetic analysis nor artificially programmed in a simple manner. Available data suggest that population dynamics and physiological control of catabolic gene expression prevail over any artificial attempt to engineer an optimal performance of the wanted catalytic activities. In this review, several valuable spin-offs of past research into genetically modified organisms with environmental applications are discussed, along with the impact of Systems Biology and Synthetic Biology in the future of environmental biotechnology.

Animals↗

COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.

SUMMARY: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. AVAILABILITY: http://maine.ebi.ac.uk:8000/services/cogent/

Computational Biology↗

Beyond 100 genomes.

By the end of 2002, we witnessed the landmark submission of the 100th complete genome sequence in the databases. An overview of these genomes reveals certain interesting trends and provides valuable insights into possible future developments.

Animals↗

Myriads of protein families, and still counting.

From the historical record of genome sequencing, we show that the rate of discovery of new families has remained constant over time, indicating that our knowledge of sequence space is far from complete.

Animals↗

Transcription regulation and environmental adaptation in bacteria.

Lifestyle can be viewed as the environment surrounding an organism and the relationships that it establishes with other species. It is one of the driving forces that contribute to the final shape of bacterial genomes. To assess how these forces affect global cellular functions, we investigated the fraction of the genome devoted to transcription-related proteins, small-molecule metabolism enzymes, and transport, for 60 bacterial genomes classified by lifestyle. Larger genomes were found to harbour more transcription factors per gene than smaller ones. In addition, free-living bacteria (with a few exceptions) are clearly enriched for transcription factors, beyond the expected proportion based on their genome size. This suggests that under complex conditions, gene expression regulation and signal integration have been strongly selected for to enable rapid adaptation to environmental conditions.

Adaptation, Biological↗

Heavy metal tolerance and metal homeostasis in Pseudomonas putida as revealed by complete genome analysis.

The genome of Pseudomonas putida KT2440 encodes an unexpected capacity to tolerate heavy metals and metalloids. The availability of the complete chromosomal sequence allowed the categorization of 61 open reading frames likely to be involved in metal tolerance or homeostasis, plus seven more possibly involved in metal resistance mechanisms. Some systems appeared to be duplicated. These might perform redundant functions or be involved in tolerance to different metals. In total, P. putida was found to bear two systems for arsenic (arsRBCH), one for chromate (chrA), four to six systems for divalent cations (two cadA and two to four czc chemiosmotic antiporters), two systems for monovalent cations: pacS, cusCBA (plus one cryptic silP gene containing a frameshift mutation), two operons for Cu chelation (copAB), one metallothionein for metal(loid) binding, one system for Te/Se methylation (tpmT) and four ABC transporters for the uptake of essential Zn, Mn, Mo and Ni (one nikABCDE, two znuACB and one mobABC). Some of the metal-related clusters are located in gene islands with atypical genome signatures. The predicted capacity of P. putida to endure exposure to heavy metals is discussed from an evolutionary perspective.

Arsenicals↗

The sigma54 regulon (sigmulon) of Pseudomonas putida.

sigma54 is unique among the bacterial sigma factors. Besides not being related in sequence with the rest of such factors, its mechanism of transcription initiation is completely different and requires the participation of a transcription activator. In addition, whereas the rest of the alternative sigma factors use to be involved in transcription of somehow related biological functions, this is not the case for sigma54 and many different and unrelated genes have been shown to be transcribed from sigma54-dependent promoters, ranging from flagellation, to utilization of several different carbon and nitrogen sources, or alginate biosynthesis. These genes have been characterized in many different bacterial species and, only until recently with the arrival of complete genome sequences, we have been able to look at the sigma54 functional role from a genomic perspective. Aided by computational methods, the sigma54 regulon has been studied both in Escherichia coli, Salmonella typhimurium and several species of the Rhizobiaceae. Here we present the analysis of the sigma54 regulon (sigmulon) in the complete genome of Pseudomonas putida KT2440. We have developed an improved method for the prediction of sigma54-dependent promoters which combines the scores of sigma54-RNAP target sequences and those of activator binding sites. In combination with other evidence obtained from the chromosomal context and the similarity with closely related bacteria, we have been able to predict more than 80% of the sigma54-dependent promoters of P. putida with high confidence. Our analysis has revealed new functions for sigma54 and, by means of comparative analysis with the previous studies, we have drawn a potential mechanism for the evolution of this regulatory system.

Bacterial Proteins↗