Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “functional annotations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗

Genomic distribution characteristics and interspecific differences of microsatellite landscapes in Felidae.

BACKGROUND: Microsatellites within genomes play crucial roles in regulating gene expression, DNA replication, and chromosomal structure and function. Analyzing the composition and distribution patterns of microsatellites in closely related species not only reveals their evolutionary dynamics and adaptive mechanisms but also provides essential technical support for applications in genetic breeding, species conservation, and disease research. As one of the world's most captivating animal groups, the landscape patterns of microsatellites across feline genomes remain to be systematically characterized. RESULTS: This study utilized high-quality genomic data to conduct a systematic comparative analysis of microsatellite landscape distribution patterns across the genomes of 13 felid species. The findings revealed that microsatellite abundance and distribution exhibit species-specific characteristics, with a non-random genomic distribution and a negative correlation between microsatellite abundance and repeat length. The predominant distribution pattern followed the sequence: single > double > quadruple > triple > quintuple > sextuple nucleotide repeats. Microsatellite abundance peaked in intergenic regions, whereas trinucleotide repeats were more prevalent within exons. Coding regions showed a marked preference for trinucleotide and hexanucleotide repeats. Enrichment analysis of GO and KEGG pathways indicated that coding sequences containing microsatellites were primarily involved in transcription and translation processes. CONCLUSIONS: Our study elucidates the distribution patterns and characteristics of microsatellites across diverse feline species, providing significant insights into their evolutionary mechanisms and functional roles. Furthermore, these findings establish a valuable reference and foundational dataset for the future development of high-quality, species-specific microsatellite markers in felids.

Animals↗

eQTM (expression quantitative trait methylation) Atlas: a comprehensive resource of over 11 million DNA methylation-gene expression associations through across 11 tissues and 4 diseases.

MOTIVATION: Epigenome-wide association studies (EWAS) have identified numerous DNA methylation (DNAm) CpG sites associated with complex traits and diseases, but interpretation of those CpG sites remains challenging because in EWAS, CpGs are mostly linked to nearby genes based only on genomic proximity. Expression quantitative trait methylation (eQTM) analyses connect DNAm CpGs with statistically associated gene expression levels. However, a comprehensive, searchable resource integrating eQTMs across diverse tissues and disease contexts has been lacking. RESULTS: We developed the eQTM Atlas, a web-based resource that manually curates more than 11 million DNAm-gene expression associations from eight cohorts, covering 11 tissue types, four broad disease contexts, 173,886 unique CpG probes and 20,231 unique genes. The Atlas supports gene- or CpG- searches by tissue or disease type and finding associated CpG or genes, visualization of cis- and trans-eQTMs through genome browser, heatmap interfaces across various tissues, and cohort-level data downloads. By integrating eQTM results with EWAS resources, the eQTM Atlas enables users to connect disease- or trait-associated CpGs to statistically associated genes rather than relying solely on proximity-based gene annotation, supporting functional interpretation of EWAS findings and generation of disease-specific regulatory hypotheses. AVAILABILITY AND IMPLEMENTATION: The eQTM Atlas is freely available at https://shiny.crc.pitt.edu/eqtm_browser/. The web interface is implemented in R Shiny and hosted through the University of Pittsburgh Center for Research Computing (CRC). Source code is available at https://github.com/ads303/eQTM-Atlas.

DNA methylation↗

[Application of DNA chip technology to biomedical research].

The completion of Human Genome Project enabled us to access to the information on nucleotide sequences of whole human genome. One of the most valuable information on human genome would be the list of approximately 35,000 genes. Although 35% of them are still needed to annotate their functions, we can genome-widely approach to various conditions including disease states. To analyze bunch of information at once, we need high-throughput technology containing most of genes. DNA chip successfully provide a stable platform technology for the massive screening of genomes. Microarrays can be used to obtain genome-wide fingerprint on transcriptional changes in various physiological and pathological conditions, leading to the mining novel genes related to those specific states. We can check the multiple molecular markers for diagnosis, prediction or prognosis of specific diseases. Data from microarray will provide huge amounts of experssion profile, which might induce the transformation of biomedical research.

Base Sequence↗

A bioinformatics-based approach for the prediction and identification of novel proteins potentially involved in phosphorylation signalling pathways.

Together with the explosion in the availability of genome data of a number of organisms including human and mouse, various methods and programs for computational prediction of protein-coding genes and annotation of functional proteins have dramatically increased. For the last decade there has been intense interest in the role of protein phosphorylation which is involved in post-translation modification mechanisms critically regulating inter/intracellular communication, patho/physiological responses and homeostasis during many biological processes. In the present study a total of 202 functionally uncharacterized human full-coding cDNA sequences were investigated using a bioinformatics-based approach. Ten novel potential substrates for protein kinases have been identified which may play multiple roles in regulating intracellular phosphorylation signalling pathways. In addition, 5 of those may be involved in the human-only post-translation mechanism regulated by specific protein kinases. The data presented here therefore would greatly contribute toward the understanding of human molecular basis and cellular signalling networks.

Amino Acid Motifs↗

Gene3D: structural assignment for whole genes and genomes using the CATH domain structure database.

We present a novel web-based resource, Gene3D, of precalculated structural assignments to gene sequences and whole genomes. This resource assigns structural domains from the CATH database to whole genes and links these to their curated functional and structural annotations within the CATH domain structure database, the functional Dictionary of Homologous Superfamilies (DHS) and PDBsum. Currently Gene3D provides annotation for 36 complete genomes (two eukaryotes, six archaea, and 28 bacteria). On average, between 30% and 40% of the genes of a given genome can be structurally annotated. Matches to structural domains are found using the profile-based method (PSI-BLAST). and a novel protocol, DRange, is used to resolve conflicts in matches involving different homologous superfamilies.

Animals↗

Re-annotating the Mycoplasma pneumoniae genome sequence: adding value, function and reading frames.

Four years after the original sequence submission, we have re-annotated the genome of Mycoplasma pneumoniae to incorporate novel data. The total number of ORFss has been increased from 677 to 688 (10 new proteins were predicted in intergenic regions, two further were newly identified by mass spectrometry and one protein ORF was dismissed) and the number of RNAs from 39 to 42 genes. For 19 of the now 35 tRNAs and for six other functional RNAs the exact genome positions were re-annotated and two new tRNA(Leu) and a small 200 nt RNA were identified. Sixteen protein reading frames were extended and eight shortened. For each ORF a consistent annotation vocabulary has been introduced. Annotation reasoning, annotation categories and comparisons to other published data on M.pneumoniae functional assignments are given. Experimental evidence includes 2-dimensional gel electrophoresis in combination with mass spectrometry as well as gene expression data from this study. Compared to the original annotation, we increased the number of proteins with predicted functional features from 349 to 458. The increase includes 36 new predictions and 73 protein assignments confirmed by the published literature. Furthermore, there are 23 reductions and 30 additions with respect to the previous annotation. mRNA expression data support transcription of 184 of the functionally unassigned reading frames.

Amino Acid Sequence↗

A draft annotation and overview of the human genome.

BACKGROUND: The recent draft assembly of the human genome provides a unified basis for describing genomic structure and function. The draft is sufficiently accurate to provide useful annotation, enabling direct observations of previously inferred biological phenomena. RESULTS: We report here a functionally annotated human gene index placed directly on the genome. The index is based on the integration of public transcript, protein, and mapping information, supplemented with computational prediction. We describe numerous global features of the genome and examine the relationship of various genetic maps with the assembly. In addition, initial sequence analysis reveals highly ordered chromosomal landscapes associated with paralogous gene clusters and distinct functional compartments. Finally, these annotation data were synthesized to produce observations of gene density and number that accord well with historical estimates. Such a global approach had previously been described only for chromosomes 21 and 22, which together account for 2.2% of the genome. CONCLUSIONS: We estimate that the genome contains 65,000-75,000 transcriptional units, with exon sequences comprising 4%. The creation of a comprehensive gene index requires the synthesis of all available computational and experimental evidence.

Chromosome Mapping↗

Protein family alignment annotation.

For bioscientists studying protein structure and function, the Protein Family Alignment Annotation Tool (Pfaat) is a useful and simple program for annotating collections of proteins. This open-source software includes methods for viewing and aligning protein families, and for annotating sequence structure and residues with known functions. It offers new options to aid the study of proteins, and an extensible annotation tool for bioinformatics developers.

Amino Acid Sequence↗

Strain-specific genes of Helicobacter pylori: distribution, function and dynamics.

Whole-genome clustering of the two available genome sequences of Helicobacter pylori strains 26695 and J99 allows the detection of 110 and 52 strain-specific genes, respectively. This set of strain-specific genes was compared with the sets obtained with other computational approaches of direct genome comparison as well as experimental data from microarray analysis. A considerable number of novel function assignments is possible using database-driven sequence annotation, although the function of the majority of the identified genes remains unknown. Using whole-genome clustering, it is also possible to detect species-specific genes by comparing the two H.pylori strains against the genome sequence of Campylobacter jejuni. It is interesting that the majority of strain-specific genes appear to be species specific. Finally, we introduce a novel approach to gene position analysis by employing measures from directional statistics. We show that although the two strains exhibit differences with respect to strain-specific gene distributions, this is due to the extensive genome rearrangements. If these are taken into account, a common pattern for the genome dynamics of the two Helicobacter strains emerges, suggestive of certain spatial constraints that may act as control mechanisms of gene flux.

Amino Acid Sequence↗

MIPS: a database for genomes and protein sequences.

The Munich Information Center for Protein Sequences (MIPS-GSF, Neuherberg, Germany) continues to provide genome-related information in a systematic way. MIPS supports both national and European sequencing and functional analysis projects, develops and maintains automatically generated and manually annotated genome-specific databases, develops systematic classification schemes for the functional annotation of protein sequences, and provides tools for the comprehensive analysis of protein sequences. This report updates the information on the yeast genome (CYGD), the Neurospora crassa genome (MNCDB), the databases for the comprehensive set of genomes (PEDANT genomes), the database of annotated human EST clusters (HIB), the database of complete cDNAs from the DHGP (German Human Genome Project), as well as the project specific databases for the GABI (Genome Analysis in Plants) and HNB (Helmholtz-Netzwerk Bioinformatik) networks. The Arabidospsis thaliana database (MATDB), the database of mitochondrial proteins (MITOP) and our contribution to the PIR International Protein Sequence Database have been described elsewhere [Schoof et al. (2002) Nucleic Acids Res., 30, 91-93; Scharfe et al. (2000) Nucleic Acids Res., 28, 155-158; Barker et al. (2001) Nucleic Acids Res., 29, 29-32]. All databases described, the protein analysis tools provided and the detailed descriptions of our projects can be accessed through the MIPS World Wide Web server (http://mips.gsf.de).

Amino Acid Sequence↗

Shared genetic architecture of obesity and gastroesophageal reflux disease.

Obesity is identified as a risk factor of gastroesophageal reflux disease (GERD). This study aims to elucidate the shared genetic architecture of obesity-related phenotypes and GERD. Based on the publicly available genome-wide association studies' datasets, this genome-wide pleiotropic association study was conducted with various genetic approaches (including linkage disequilibrium score regression, high-definition likelihood inference for genetic correlations, pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis) sequentially to unravel the genetic associations from single-nucleotide polymorphism to gene levels, and to reveal the underlying shared genetic architecture between obesity-related phenotypes and GERD. This study discovered shared genetic mechanisms between GERD and several obesity-related phenotypes, including arm fat percentage (left), arm fat percentage (right), leg fat percentage (left), leg fat percentage (right), trunk fat percentage, waist-to-hip ratio, and body mass index. Significant genetic correlations were observed by linkage disequilibrium score regression and high-definition likelihood inference for genetic correlations, with multiple associated pleiotropic loci and their mapped genes identified by pleiotropic analysis under composite null hypothesis, Functional Mapping and Annotation, Bayesian colocalization, summary-based Mendelian randomization, and multi-marker analysis of genomic annotation analysis. Additionally, several brain tissues were identified to be linked to both obesity and GERD by multi-marker analysis of genomic annotation. This research provided strong evidence of genetic correlations and brought novel insights into the underlying genetic connections and shared genetic architectures of obesity and GERD.

Humans↗

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics↗

Intrinsic errors in genome annotation.

Genome sequencing is usually followed by routine annotation of protein function based on the assumption that similar sequences will have similar functions. Here, we introduce a simple calculation to estimate the magnitude of any possible annotation errors. We counted the number of discrepancies in the annotation of well-established sets of similar proteins and extrapolated these values to the pairs of similar sequences used for the annotation of different microbial genomes. We conclude that the number of potential errors in the prediction of detailed functions is higher than is usually believed.

Binding Sites↗

Arabidopsis genes involved in acyl lipid metabolism. A 2003 census of the candidates, a study of the distribution of expressed sequence tags in organs, and a web-based database.

The genome of Arabidopsis has been searched for sequences of genes involved in acyl lipid metabolism. Over 600 encoded proteins have been identified, cataloged, and classified according to predicted function, subcellular location, and alternative splicing. At least one-third of these proteins were previously annotated as "unknown function" or with functions unrelated to acyl lipid metabolism; therefore, this study has improved the annotation of over 200 genes. In particular, annotation of the lipolytic enzyme group (at least 110 members total) has been improved by the critical examination of the biochemical literature and the sequences of the numerous proteins annotated as "lipases." In addition, expressed sequence tag (EST) data have been surveyed, and more than 3,700 ESTs associated with the genes were cataloged. Statistical analysis of the number of ESTs associated with specific cDNA libraries has allowed calculation of probabilities of differential expression between different organs. More than 130 genes have been identified with a statistical probability > 0.95 of preferential expression in seed, leaf, root, or flower. All the data are available as a Web-based database, the Arabidopsis Lipid Gene database (http://www.plantbiology.msu.edu/lipids/genesurvey/index.htm). The combination of the data of the Lipid Gene Catalog and the EST analysis can be used to gain insights into differential expression of gene family members and sets of pathway-specific genes, which in turn will guide studies to understand specific functions of individual genes.

Acylation↗

The Lipase Engineering Database: a navigation and analysis tool for protein families.

The Lipase Engineering Database (LED) (http://www.led.uni-stuttgart.de) integrates information on sequence, structure, and function of lipases, esterases, and related proteins. Sequence data on 806 protein entries are assigned to 38 homologous families, which are grouped into 16 superfamilies with no global sequence similarity between each other. For each family, multisequence alignments are provided with functionally relevant residues annotated. Pre-calculated phylogenetic trees allow navigation inside superfamilies. Experimental structures of 45 proteins are superposed and consistently annotated. The LED has been applied to systematically analyze sequence-structure-function relationships of this vast and diverse enzyme class. It is a useful tool to identify functionally relevant residues apart from the active site residues, and to design mutants with desired substrate specificity.

Amino Acid Sequence↗

Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.

Functional characterizations of thousands of gene products from many species are described in the published literature. These discussions are extremely valuable for characterizing the functions not only of these gene products, but also of their homologs in other organisms. The Gene Ontology (GO) is an effort to create a controlled terminology for labeling gene functions in a more precise, reliable, computer-readable manner. Currently, the best annotations of gene function with the GO are performed by highly trained biologists who read the literature and select appropriate codes. In this study, we explored the possibility that statistical natural language processing techniques can be used to assign GO codes. We compared three document classification methods (maximum entropy modeling, naïve Bayes classification, and nearest-neighbor classification) to the problem of associating a set of GO codes (for biological process) to literature abstracts and thus to the genes associated with the abstracts. We showed that maximum entropy modeling outperforms the other methods and achieves an accuracy of 72% when ascertaining the function discussed within an abstract. The maximum entropy method provides confidence measures that correlate well with performance. We conclude that statistical methods may be used to assign GO codes and may be useful for the difficult task of reassignment as terminology standards evolve over time.

Algorithms↗

The CATH Dictionary of Homologous Superfamilies (DHS): a consensus approach for identifying distant structural homologues.

A consensus approach has been developed for identifying distant structural homologues. This is based on the CATH Dictionary of Homologous Superfamilies (DHS), a database of validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies (URL: http://www. biochem.ucl.ac.uk/bsm/dhs). Multiple structural alignments have been generated for 362 well-populated superfamilies in the CATH structural domain database and annotated with secondary structure, physicochemical properties, functional sequence patterns and protein-ligand interaction data. Consensus functional information for each superfamily includes descriptions and keywords extracted from SWISS-PROT and the ENZYME database. The Dictionary provides a powerful resource to validate, examine and visualize key structural and functional features of each homologous superfamily. The value of the DHS, for assessing functional variability and identifying distant evolutionary relationships, is illustrated using the pyridoxal-5'-phosphate (PLP) binding aspartate aminotransferase superfamily. The DHS also provides a tool for examining sequence-structure relationships for proteins within each fold group.

Amino Acid Sequence↗