Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

The Complete Chloroplast Genome and the Phylogenetic Analysis of Panicum bisulcatum (Thumb.) (Poaceae).

The chloroplast (cp) genome of Panicum bisulcatum (Thumb.), a significant agricultural weed, was sequenced and characterized to elucidate its genomic architecture, evolutionary dynamics, and phylogenetic relationships. The complete cp genome was assembled as a circular DNA molecule of 138,489 bp, exhibiting a typical quadripartite structure comprising a large single-copy (LSC, 82,260 bp), a small single-copy (SSC, 12,569 bp), and a pair of inverted repeats (IR, 21,830 bp each) regions. It encodes 135 genes, including 89 protein-coding genes, 49 tRNAs, and 8 rRNAs. Functional annotation revealed that most genes are involved in photosynthesis and genetic system. A total of 51 simple sequence repeats (SSRs) and 62 long repeats (LRs) were identified, providing potential molecular markers. Comparative analysis of IR boundaries highlighted both conserved features and species-specific expansion/contraction events among Panicum species. Phylogenomic analysis robustly placed P. bisulcatum within the genus Panicum, showing a closest relationship with P. incomtum and confirming the monophyly of the genus. Furthermore, single nucleotide polymorphism (SNP) analysis with its closest relative, P. incomtum, revealed 4659 SNPs, with a dominance of synonymous substitutions, indicating the action of purifying selection. This study provides the first comprehensive cp genomic resource for P. bisulcatum, which will facilitate future studies in species identification, phylogenetic reconstruction, population genetics, and the development of sustainable management strategies for this weed.

Phylogeny↗

Deciphering the Genetic Underpinnings of Liver Cirrhosis-Heart Failure Comorbidity Through Multi-Omics: CRIM1 as a Key Endothelial Mediator.

The co-occurrence of liver cirrhosis (LC) and heart failure (HF) poses considerable clinical challenges, yet the cellular and molecular determinants of this comorbidity remain poorly characterized. To address this, we developed an integrative multi-omics pipeline encompassing GWAS meta-analysis, gsMap-based spatial transcriptomic projection, GeneEnrich functional annotation, single-cell atlas construction, seismicGWAS and ECLIPSER cell-type scoring, eCAVIAR and fastenloc colocalization, hdWGCNA network inference, scTenifoldKnk in silico gene perturbation, and GCTA-COJO fine-mapping. Quality-controlled meta-analysis yielded 12,347,758 and 9,256,862 variant-level associations for LC and HF, respectively. Spatial projection confirmed preferential enrichment of disease signals within embryonic hepatic and cardiac compartments. Pathway analyses disclosed that LC-linked loci were concentrated in lipid metabolic programs, whereas HF-linked loci implicated mitochondrial bioenergetics and lysosomal degradation. At the cellular level, endothelial cells emerged as the dominant HF-associated population. Convergent evidence from five orthogonal algorithms pinpointed CRIM1 as the sole robustly supported shared gene, selectively enriched in HF endothelial cells; virtual perturbation further identified LCP1 and PTPRC as downstream regulatory nodes. Fine-mapping of the chromosome 2 locus harboring rs12476437 revealed multiple statistically independent signals in the vicinity of CRIM1. Collectively, these findings computationally prioritize the endothelial-CRIM1 axis as a previously unappreciated candidate mechanistic bridge between LC and HF requiring experimental validation.

Humans↗

A Network Pharmacology and Molecular Docking Study of TongBi Formula for Osteoarthritis.

This study applied network pharmacology combined with molecular docking to predict the potential therapeutic targets and molecular mechanisms of TongBi Formula (TBF) in osteoarthritis (OA). Active components and corresponding targets of TBF were retrieved from the traditional Chinese medicine Systems Pharmacology Database and Analysis Platform, while OA-related targets were collected from Online Mendelian Inheritance in Man, GeneCards, DrugBank, and Therapeutic Target Database. A network visualization and analysis software was used to construct compound-target and protein-protein interaction (PPI) networks. Gene Ontology functional annotation and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed using the Database for Annotation, Visualization and Integrated Discovery platform. Molecular docking analysis was conducted using a molecular docking software to evaluate the predicted binding affinity between key active compounds and core target proteins. A total of 47 overlapping targets between TBF and OA were identified. PPI network analysis highlighted JUN, RELA, IL6, MAPK1, and IL10 as potential hub targets. Enrichment analysis suggested that TBF may regulate inflammation, lipid metabolism, and multiple intracellular signaling pathways associated with OA progression. Molecular docking results demonstrated favorable predicted binding affinities between core active compounds and key OA-related protein targets. These findings provide a computational framework for understanding the potential mechanisms of TBF against OA and support further experimental validation.

Molecular Docking Simulation↗

A comparative genomic analysis of left- and right-sided colon cancer using real-world data from the AACR project GENIE BPC dataset.

Left- and Right-sided colon cancers (LCC and RCC) are increasingly recognized as distinct clinicopathological and molecular subtypes with divergent prognoses and therapeutic responses. Leveraging a large, multi-institutional cohort from the AACR Project Genomics Evidence Neoplasia Information Exchange (GENIE) Biopharma Collaborative (BPC) (n = 750; LCC: 363 vs. RCC: 387), we conducted a comprehensive analysis of mutational profiles, tumor mutation burden (TMB), and survival outcomes. Our findings revealed a markedly higher TMB in RCC compared to LCC (6.65 &#xb1; 11.3 vs. 3.17 &#xb1; 4.35; adjusted P = 3.12&#xd7;10-32), suggesting greater genomic instability in RCC. After applying functional annotation filters (PolyPhen > 0.85, SIFT < 0.05), RCC tumors were significantly enriched for mutations in BRAF (23.1% vs. 6.7%), KMT2D (8.6% vs. 3.2%), and SMAD4 (13.1% vs. 7.3%), while TP53 mutations predominated in LCC (40.6% vs. 31.8%). Multivariate Cox regression analysis identified RCC as an independent predictor of poorer overall survival (OS) relative to LCC (HR: 1.30, 95% CI: 1.02-1.66, P = 0.033). Notably, KRAS mutations were associated with significantly worse OS in LCC (HR: 1.68, 95% CI: 1.06-2.70, P = 0.027), while BRAF mutations predicted adverse outcomes in RCC (HR: 1.58, 95% CI: 1.05-2.37, P = 0.028). These results underscore the prognostic value of tumor sidedness and specific genetic alterations in colon adenocarcinoma. Our study highlights the need for sidedness-specific molecular profiling to inform precision oncology strategies in colon cancer management.

BRAF↗

Multi-Ancestry Survival GWAS of Substance Use Initiation in the ABCD Study.

BACKGROUND: Substance use initiation in adolescence is influenced by both genetic and environmental factors; however, large-scale genetic studies often treat initiation as a binary outcome and underuse longitudinal timing information. METHODS: We conducted time-to-event (survival) genome-wide association analyses (GWAS) of initiation for four outcomes-alcohol, nicotine, cannabis, and any substance use-using longitudinal follow-up data from the Adolescent Brain Cognitive Development (ABCD) Study. We performed ancestry-stratified GWAS within European (EUR), African (AFR), and Hispanic (HISP) groups, applying consistent quality control and covariate adjustment. Summary statistics were harmonized across ancestries and meta-analyzed using inverse-variance weighted fixed-effects and DerSimonian-Laird random-effects models. We evaluated genomic inflation and heterogeneity (Cochran's Q and I 2), identified independent lead variants at genome-wide and suggestive significance thresholds, and assessed cross-trait overlap of associated loci. RESULTS: In the multi-ancestry meta-analysis, we observed suggestive association signals across traits (minimum p-values: alcohol ~ 1 &#xd7; 10-7, any ~ 1 &#xd7; 10-7, cannabis ~ 5 &#xd7; 10-8, nicotine ~ 1 &#xd7; 10-8). Nicotine initiation showed one genome-wide significant variant in both fixed- and random-effects meta-analyses (p < 5 &#xd7; 10-8). Across traits, suggestive loci demonstrated limited overlap, with the strongest concordance between alcohol and any substance use, consistent with shared liability. Heterogeneity statistics indicated that some loci exhibited cross-ancestry variation in effect estimates. CONCLUSIONS: Survival GWAS leveraging initiation timing can identify genetic signals that may be missed by binary designs and enables principled multi-ancestry synthesis. Our results highlight both shared and trait-specific genetic contributions to early substance initiation and provide a foundation for downstream functional annotation and integrative modeling with environmental risk factors. These findings demonstrate the value of incorporating developmental timing into genetic discovery and provide a framework for integrating longitudinal risk modeling with genomic analyses.

ABCD↗

Building dictionaries of 1D and 3D motifs by mining the Unaligned 1D sequences of 17 archaeal and bacterial genomes.

We have used the Teiresias algorithm to carry out unsupervised pattern discovery in a database containing the unaligned ORFs from the 17 publicly available complete archaeal and bacterial genomes and build a 1D dictionary of motifs. These motifs which we refer to as seqlets account for and cover 97.88% of this genomic input at the level of amino acid positions. Each of the seqlets in this 1D dictionary was located among the sequences in Release 38.0 of the Protein Data Bank and the structural fragments corresponding to each seqlet's instances were identified and aligned in three dimensions: those of the seqlets that resulted in RMSD errors below a pre-selected threshold of 2.5 Angstroms were entered in a 3D dictionary of structurally conserved seqlets. These two dictionaries can be thought of as cross-indices that facilitate the tackling of tasks such as automated functional annotation of genomic sequences, local homology identification, local structure characterization, comparative genomics, etc.

Algorithms↗

A probabilistic learning approach to whole-genome operon prediction.

We present a computational approach to predicting operons in the genomes of prokaryotic organisms. Our approach uses machine learning methods to induce predictive models for this task from a rich variety of data types including sequence data, gene expression data, and functional annotations associated with genes. We use multiple learned models that individually predict promoters, terminators and operons themselves. A key part of our approach is a dynamic programming method that uses our predictions to map every known and putative gene in a given genome into its most probable operon. We evaluate our approach using data from the E. coli K-12 genome.

Gene Expression Profiling↗

Bioinformatics of large-scale protein interaction networks.

We survey recent techniques for construction and prediction of large-scale protein interaction networks, focusing on computational processing steps. Special emphasis is placed on critical assessment of data completeness and reliability of the various approaches. Once built, protein interaction networks can be used for functional annotation or to generate higher-level biological hypotheses on pathways.

Bacterial Proteins↗

Global analysis of large-scale chemical and biological experiments.

Research in the life sciences is increasingly dominated by high-throughput data collection methods that benefit from a global approach to data analysis. Recent innovations that facilitate such comprehensive analyses are highlighted. Several developments enable the study of the relationships between newly derived experimental information, such as biological activity in chemical screens or gene expression studies, and prior information, such as physical descriptors for small molecules or functional annotation for genes. The way in which global analyses can be applied to both chemical screens and transcription profiling experiments using a set of common machine learning tools is discussed.

Animals↗

Peptomics, identification of novel cationic Arabidopsis peptides with conserved sequence motifs.

Few plant peptides involved in intercellular communication have been experimentally isolated. Sequence analysis of the Arabidopsis thaliana genome has revealed numerous transmembrane receptors predicted to bind proteinacious ligands, emphasizing the importance of identifying peptides with signaling function. Annotation of the Arabidopsis genome sequence has made it possible to identify peptide-encoding genes. However, such annotational identification is impeded because small genes are poorly predicted by gene-prediction algorithms, thus prompting the alternative approaches described here. We initially performed a systematic analysis of short polypeptides encoded by annotated genes on two Arabidopsis chromosomes using SignalP to identify potentially secreted peptides. Subsequent homology searches with selected, putatively secreted peptides, led to the identification of a potential, large Arabidopsis family of 34 genes. The predicted peptides are characterized by a conserved C-terminal sequence motif and additional primary structure conservation in a core region. The majority of these genes had not previously been annotated. A subset of the predicted peptides show high overall sequence similarity to Rapid Alkalinization Factor (RALF), a peptide isolated from tobacco. We therefore refer to this peptide family as RALFL for RALF-Like. RT-PCR analysis confirmed that several of the Arabidopsis genes are expressed and that their expression patterns vary. The identification of a large gene family in the genome of the model organism Arabidopsis thaliana demonstrates that a combination of systematic analysis and homology searching can contribute to peptide discovery.

Algorithms↗

Differential responses of stress genes to low dose-rate gamma irradiation.

In the past, most mechanistic studies of ionizing radiation response have employed very large doses, then extrapolated the results down to doses relevant to human exposure. It is becoming increasingly apparent, however, that this does not give an accurate or complete picture of the effects of most environmental exposures, which tend to be of low dose and protracted over time. We have initiated direct studies of low dose exposures, and using the relatively responsive ML-1 cell line, have shown that changes in gene expression can be triggered by doses of gamma-rays of 10 cGy and less in human cells. We have now extended these studies to investigate the effects on gene induction of reducing the rate of irradiation. In the ML-1 human myeloid leukemia cell line, we have found that reducing the dose rate over three orders of magnitude results in some protection against the induction of apoptosis, but still causes linear induction of the p53-regulated genes CDKN1A, GADD45A, and MDM2 between 2 and 50 cGy. Reducing the rate of exposure reduces the magnitude of induction of CDKN1A and GADD45A, but not the magnitude or duration of cell cycle delay. In contrast, MDM2 is induced to the same extent regardless of the rate of dose delivery. Microarray analysis has identified additional low dose-rate-inducible genes, and indicates the existence of two general classes of low dose-rate responders in ML-1. One group of genes is induced in a dose rate-dependent fashion, similar to GADD45A and CDKN1A. Functional annotation of this gene cluster indicates a preponderance of genes with known roles in apoptosis regulation. Similarly, a group of genes with dose rate-independent induction, such as seen for MDM2, was also identified. The majority of genes in this group are involved in cell cycle regulation. This apparent differential regulation of stress signaling pathways and outcomes in response to protracted radiation exposure has implications for carcinogenesis and risk assessment, and could not have been predicted from classical high dose studies.

Apoptosis↗

Accelerating comparative genomics using parallel computing.

In the past decade there has been an increase in the number of completely sequenced genomes due to the race of multibillion-dollar genome-sequencing projects. The enormous biological sequence data thus flooding into the sequence databases necessitates the development of efficient tools for comparative genome sequence analysis. The information deduced by such analysis has various applications viz. structural and functional annotation of novel genes and proteins, finding gene order in the genome, gene fusion studies, constructing metabolic pathways etc. Such study also proves invaluable for pharmaceutical industries, such as in silico drug target identification and new drug discovery. There are various sequence analysis tools available for mining such useful information of which FASTA and Smith-Waterman algorithms are widely used. However, analyzing large datasets of genome sequences using the above codes seems to be impractical on uniprocessor machines. Hence there is a need for improving the performance of the above popular sequence analysis tools on parallel cluster computers. Performance of the Smith-Waterman (SSEARCH) and FASTA programs were studied on PARAM 10000, a parallel cluster of workstations designed and developed in-house. FASTA and SSEARCH programs, which are available from the University of Virginia, were ported on PARAM and were optimized. In this era of high performance computing, where the paradigm is shifting from conventional supercomputers to the cost-effective general-purpose cluster of workstations and PCs, this study finds extreme relevance. Good performance of sequence analysis tools on a cluster of workstations was demonstrated, which is important for accelerating identification of novel genes and drug targets by screening large databases.

Algorithms↗

Immunoinformatics--the new kid in town.

The astounding diversity of immune system components (e.g. immunoglobulins, lymphocyte receptors, or cytokines) together with the complexity of the regulatory pathways and network-type interactions makes im munology a combinatorial science. Currently available data represent only a tiny fraction of possible situations and data continues to accrue at an exponential rate. Computational analysis has therefore become an essential element of immunology research with a main role of immunoinformatics being the management and analysis of immunological data. More advanced analyses of the immune system using computational models typically involve conversion of an immunological question to a computational problem, followed by solving of the computational problem and translation of these results into biologically meaningful answers. Major immunoinformatics developments include immunological databases, sequence analysis, structure modelling, mathematical modelling of the immune system, simulation of laboratory experiments, statistical support for immunological experimentation and immunogenomics. In this paper we describe the status and challenges within these sub-fields. We foresee the emergence of immunomics not only as a collective endeavour by researchers to decipher the sequences of T cell receptors, immunoglobulins, and other immune receptors, but also to functionally annotate the capacity of the immune system to interact with the whole array of selfand non-self entities, including genome-to-genome interactions.

Allergy and Immunology↗

Bioinformatics of the Paracoccidioides brasiliensis EST Project.

Paracoccidioides brasiliensis is the etiological agent of paracoccidioidomycosis, an endemic mycosis of Latin America. This fungus presents a dimorphic character; it grows as a mycelium at room temperature, but it is isolated as yeast from infected individuals. It is believed that the transition from mycelium to yeast is important for the infective process. The Functional and Differential Genome of Paracoccidioides brasiliensis Project--PbGenome Project was developed to study the infection process by analyzing expressed sequence tags--ESTs, isolated from both mycelial and yeast forms. The PbGenome Project was executed by a consortium that included 70 researchers (professors and students) from two sequencing laboratories of the midwest region of Brazil; this project produced 25,741 ESTs, 19,718 of which with sufficient quality to be analyzed. We describe the computational procedures used to receive process, analyze these ESTs, and help with their functional annotations; we also detail the services that were used for sequence data exploration. Various programs were compared for filtering and grouping the sequences, and they were adapted to a user-friendly interface. This system made the analysis of the differential transcriptome of P. brasiliensis possible.

Brazil↗

Deciphering protein network organization using phylogenetic profile groups.

Phylogenetic profiling is now an effective computational method to detect functional associations between proteins. The method links two proteins in accordance with the similarity of their phyletic distributions across a set of genomes. While pair-wise linkage is useful, it misses correlations in higher order groups: triplets, quadruplets, and so on. Here we assess the probability of observing co-occurrence patterns of 3 binary profiles by chance and show that this probability is asymptotically the same as the mutual information in three profiles. We demonstrate the utility of the probability and the mutual information metrics in detecting overly represented triplets of orthologous proteins which could not be detected using pairwise profiles. These triplets serve as small building blocks, i.e. motifs in protein networks; they allow us to infer the function of uncharacterized members, and facilitate analysis of the local structure and global organization of the protein network. Our method is extendable to N-component clusters, and therefore serves as a general tool for high order protein function annotation.

Amino Acid Motifs↗

Gene Class expression: analysis tool of Gene Ontology terms with gene expression data.

Serial analysis of gene expression (SAGE) technology produces large sets of interesting genes that are difficult to analyze directly. Bioinformatics tools are needed to interpret the functional information in these gene sets. We present an interactive web-based tool, called Gene Class, which allows functional annotation of SAGE data using the Gene Ontology (GO) database. This tool performs searches in the GO database for each SAGE tag, making associations in the selected GO category for a level selected in the hierarchy. This system provides user-friendly data navigation and visualization for mapping SAGE data onto the gene ontology structure. This tool also provides graphical visualization of the percentage of SAGE tags in each GO category, along with confidence intervals and hypothesis testing.

Animals↗

Identification of differentially expressed genes in human bladder cancer through genome-wide gene expression profiling.

Large-scale gene expression profiling is an effective strategy for understanding the progression of bladder cancer (BC). The aim of this study was to identify genes that are expressed differently in the course of BC progression and to establish new biomarkers for BC. Specimens from 21 patients with pathologically confirmed superficial (n = 10) or invasive (n = 11) BC and 4 normal bladder samples were studied; samples from 14 of the 21 BC samples were subjected to microarray analysis. The validity of the microarray results was verified by real-time RT-PCR. Of the 136 up-regulated genes we detected, 21 were present in all 14 BCs examined (100%), 44 in 13 (92.9%), and the other 71 in 12 BCs (85.7%). Of 69 down-regulated genes, 25 were found in all 14 BCs (100%), 22 in 13 (92.9%), and the other 22 in 12 BCs (85.7%). Functional annotation revealed that of the up-regulated genes, 36% were involved in metabolism and 14% in transcription and processing; 25% of the down-regulated genes were linked to cell adhesion/surface and 21% to cytoskeleton/cell membrane. Real-time RT-PCR confirmed the microarray results obtained for the 6 most highly up- and the 2 most highly down-regulated genes. Among the 6 most highly up-regulated genes, CKS2 was the only gene with a significantly greater level of up-regulation in invasive than in superficial BC (p = 0.04). To confirm this result, we subjected all 21 BC samples to real-time PCR assay for CKS2. We found a considerable difference between superficial and invasive BC (p = 0.001). Interestingly, there was a considerable difference between the normal bladder and invasive BC (p = 0.001) and less difference between the normal bladder and superficial BC (p = 0.005). We identified several genes as promising candidates for diagnostic biomarkers of human BC and the CKS2 gene not only as a potential biomarker for diagnosing, but also for staging human BC. This is the first report demonstrating that CKS2 expression is strongly correlated with the progression of human BC.

Aged↗

Holter recordings with continuous marker annotation to evaluate pacemaker function.

BACKGROUND: Pacemaker marker annotations facilitate the interpretation of device behavior in addition to ECG recordings. However, they are only available in conjunction with a programmer. We studied the diagnostic value of a prototype Telemetry Holter Decoder (THD), providing continuous marker annotations on a conventional Holter. METHODS: The study included 20 patients with VDD or DDDR pacemakers. A 24-hour Holter was performed using the THD. Marker annotations are transmitted from the pacemaker to the THD, which transforms them into analog signals, which are recorded on one of the Holter channels. RESULTS: During a total recording time of 458 hours, high quality marker annotations were retrieved for every patient. Artefacts disturbed the recordings during 184 min (0.67%). The THD provided information not discernible on the ECG: intermittent atrial undersensing during sinus rhythm (1096 times). Atrial tachycardias, not visible on the ECG, were detected in 2 patients. The activation of tachycardia response algorithms was clearly annotated in 11,516 events. A total of 8875 PVC's occurred, 57.8% of which were classified incorrectly in the event counters as conducted or fusion beats. Atrial far-field sensing or VA conduction was demonstrated 4294 times. Electromagnetic interferences, not visible on the ECG, could be seen three times. CONCLUSION: Recording of continuous high-quality marker annotation on a conventional Holter is feasible. The THD provides important information on device behavior, even in patients assumed to have regular device function, and shows to be clearly superior to ECG interpretation alone. Such data can be used for improved programming, troubleshooting and for the validation of new algorithms.

Aged↗