Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Systematic genome-wide approach to positional candidate cloning for identification of novel human disease genes.

BACKGROUND: Recent large-scale genome projects afford a unique opportunity to identify many novel disease genes and thereby better understand the genetic basis of human disease. Functional Annotation of Mouse (FANTOM) 2, the largest mouse transcriptome project yet, provides a wealth of data on novel genes, splice variants and non-coding RNA, and provides a unique opportunity to identify novel human disease genes. AIMS: To demonstrate the power of combining the FANTOM 2 cDNA dataset with a positional candidate approach and bioinformatics analysis to identify genes underlying human genetic disease. RESULTS: By mapping all FANTOM 2 cDNA to the human genome, we were able to identify mouse clones that co-localised on the human genome with mapped but uncloned human disease loci. By this method we identified mouse and corresponding human genes mapping within the loci of 100 different human genetic diseases (mapped interval of <5 cM). Of particular interest was the elucidation through FANTOM 2 novel mouse gene data of candidate human genes for the following: (i) developmental -disorders: neural tube defect, Meckel syndrome, Wolf--Hirschhorn syndrome and keratosis follicularis spinulosa decalvans cum ophiasi; (ii) neurological disorders: benign familial infantile convulsions 3, early-onset cerebellar ataxia with retained tendon reflexes, infantile-onset spinocerebellar ataxia and vacuolar neuro-myopathy and (iii) cancer-related syndromes: tylosis with oesophageal cancer and low-grade B-cell chronic lymphatic leukaemia. CONCLUSIONS: The FANTOM 2 data will dramatically accelerate efforts to identify genes underlying human disease. It will also facilitate the creation of transgenic mouse models to help elucidate the function of potential human disease genes.

Animals↗

Histone Sequence Database: sequences, structures, post-translational modifications and genetic loci.

The Histone Sequence Database is an annotated and searchable collection of all available histone and histone fold sequences and structures. Particular emphasis has been placed on documenting conflicts between similar sequence entries from a number of source databases, conflicts that are not necessarily documented in the source databases themselves. New additions to the database include compilations of post-translational modifications for each of the core and linker histones, as well as genomic information in the form of map loci for the human histone gene complement, with the genetic loci linked to Online Mendelian Inheritance in Man (OMIM). The database is freely accessible through the World Wide Web at either http://genome.nhgri.nih.gov/histones/ or http://www.ncbi.nlm.nih. gov/Baxevani/HISTONES

Animals↗

Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project.

BackgroundPrior studies examined variants within presenilin-2 (PSEN2), presenilin-1 (PSEN1), and amyloid precursor protein (APP) genes. However, previously-reported clinically-relevant variants and other predicted damaging missense (DM) variants have not been characterized in a newer release of the Alzheimer's Disease Sequencing Project (ADSP).ObjectiveTo characterize previously-reported clinically-relevant variants and DM variants in PSEN2, PSEN1, APP within the participants from the ADSP.MethodsWe identified rare variants (MAF&#x2009;<&#x2009;1%) in PSEN2, PSEN1, and APP in 14,641 individuals with whole genome sequencing and 16,849 individuals with whole exome sequencing available (Ntotal&#x2009;=&#x2009;31,490). We additionally curated variants from ClinVar, OMIM, and Alzforum and report carriers of variants in clinical databases as well as predicted DM variants in these genes.ResultsWe detected 31 previously-reported clinically-relevant variants with alternate alleles observed within the ADSP: 4 variants in PSEN2, 25 in PSEN1, and 2 in APP. The overall variant carrier rate for the 31 clinically-relevant variants in the ADSP was 0.3%. We observed that 79.5% of the variant carriers were cases compared to 3.9% were controls. In those with AD, the mean age of onset of AD among carriers of these clinically-relevant variants was 19.6&#x2009;&#xb1;&#x2009;1.4 years earlier compared with noncarriers (p&#x2009;=&#x2009;7.8&#x2009;&#xd7;&#x2009;10-57). Additionally, we identified 197 rare variants (MAF&#x2009;<&#x2009;1%) within ADSP participants not reported in known clinical databases.ConclusionsA small proportion of individuals in the ADSP are carriers of a previously-reported clinically-relevant variant allele for AD and these participants have significantly earlier age of AD onset compared to noncarriers.

Humans↗

Prediction of protein function from protein sequence and structure.

The sequence of a genome contains the plans of the possible life of an organism, but implementation of genetic information depends on the functions of the proteins and nucleic acids that it encodes. Many individual proteins of known sequence and structure present challenges to the understanding of their function. In particular, a number of genes responsible for diseases have been identified but their specific functions are unknown. Whole-genome sequencing projects are a major source of proteins of unknown function. Annotation of a genome involves assignment of functions to gene products, in most cases on the basis of amino-acid sequence alone. 3D structure can aid the assignment of function, motivating the challenge of structural genomics projects to make structural information available for novel uncharacterized proteins. Structure-based identification of homologues often succeeds where sequence-alone-based methods fail, because in many cases evolution retains the folding pattern long after sequence similarity becomes undetectable. Nevertheless, prediction of protein function from sequence and structure is a difficult problem, because homologous proteins often have different functions. Many methods of function prediction rely on identifying similarity in sequence and/or structure between a protein of unknown function and one or more well-understood proteins. Alternative methods include inferring conservation patterns in members of a functionally uncharacterized family for which many sequences and structures are known. However, these inferences are tenuous. Such methods provide reasonable guesses at function, but are far from foolproof. It is therefore fortunate that the development of whole-organism approaches and comparative genomics permits other approaches to function prediction when the data are available. These include the use of protein-protein interaction patterns, and correlations between occurrences of related proteins in different organisms, as indicators of functional properties. Even if it is possible to ascribe a particular function to a gene product, the protein may have multiple functions. A fundamental problem is that function is in many cases an ill-defined concept. In this article we review the state of the art in function prediction and describe some of the underlying difficulties and successes.

Amino Acid Sequence↗

Characterizing trends in clinical genetic testing: A single-center analysis of EHR data from 1.8 million patients over two decades.

A lack of structural data in electronic health records (EHRs) makes assessing the impact of genetic testing on clinical practice challenging. We extracted clinical genetic tests from the EHRs of more than 1.8 million patients seen at Vanderbilt University Medical Center from 2002 to 2022. With these data, we quantified the use of clinical genetic testing in healthcare and described how testing patterns and results changed over time. We assessed trends in types of genetic tests, tracked usage across medical specialties, and introduced a new measure, the genetically attributable fraction (GAF), to quantify the proportion of observed phenotypes attributable to a genetic diagnosis over time. We identified 104,392 tests and 19,032 molecularly confirmed diagnoses. The proportion of patients with genetic testing in their EHRs increased from 1.0% in 2002 to 6.1% in 2022, and testing became more comprehensive with the growing use of multi-gene panels. The number of unique diseases diagnosed with genetic testing increased from 51 in 2002 to 509 in 2022, and there was a rise in the number of variants of uncertain significance. The phenome-wide GAF for 6,505,620 diagnoses made in 2022 was 0.46%, and the GAF was greater than 5% for 74 phenotypes, including pancreatic insufficiency (67%), chorea (64%), atrial septal defect (24%), microcephaly (17%), paraganglioma (17%), and ovarian cancer (6.8%). Our study provides a comprehensive quantification of the increasing role of genetic testing at a major academic medical institution and demonstrates its growing utility in explaining the observed medical phenome.

Humans↗

Expression profiling of uniparental mouse embryos is inefficient in identifying novel imprinted genes.

Imprinted genes are expressed from only one allele in a parent-of-origin-specific manner. We here describe a systematic approach to identify novel imprinted genes using quantification of allele-specific expression by Pyrosequencing, a highly accurate method to detect allele-specific expression differences. Sixty-eight candidate imprinted transcripts mapping to known imprinted chromosomal regions were selected from a recent expression profiling study of uniparental mouse embryos and analyzed. Three novel imprinted transcripts encoding putative non-protein-coding RNAs were identified on the basis of parent-of-origin-specific monoallelic expression in E11.5 (C57BL/6 x Cast/Ei)F1 and informative (C57BL/6 x Cast/Ei) x C57BL/6 backcross embryos. In addition, four transcripts with preferential expression of a strain-specific allele were found. Intriguingly, a vast majority of the analyzed transcripts showed no imprinting-associated expression in F1 embryos. These data strengthen the view that a large fraction of nonimprinted genes is differentially expressed between parthenogenetic and androgenetic embryos and question the efficiency of expression profiling of uniparental embryos to identify novel imprinted genes.

Alleles↗

Shared susceptibility region on chromosome 15 between autism and catatonia.

We have compiled significant linkage results from 20 genome scans for the autism syndrome disorder (ASD) and 2 for catatonia in schizophrenia (SZ). Localization of the markers has been updated across the studies using the same cytological (Genetic Location Database), physical (National Center for Biological Information), and genetic (Marshfield) maps. Eight autosomal chromosomes (1, 2, 3, 7, 9, 13, 15, and 17) showed significant linkages with ASD, and one with catatonia (15). Chromosome 15 was further characterized for SZ genome scans (N = 4) since catatonia was observed in SZ patients, for candidate genes for ASD and catatonia, and for the numerous chromosomal rearrangement and abnormalities associated to ASD. From these results, we observed that four potential susceptibility regions for ASD could be observed on chromosome 15 at 15q11-q13, 15q14-q21, 15q22-q23, and 15q26, respectively. All the four regions were shared between ASD and SZ, with 15q15-q21 being also shared with catatonia. Strong candidate genes, such as gamma-aminobutyric acid receptor B3, A5, and G3, have shown associations with ASD at 15q11-q13 susceptibility region where the majority of the chromosomal rearrangements are also found. On the other hand, negative association results were observed at 15q14-q21 susceptibility region for catatonia with the genes encoding the zinc transporter SLC30A4, the cholinergic receptor nicotinic alpha polypeptide 7, and the delta-like 4 Drosophila. Further, fine mapping and candidate gene analyses are needed to highlight potential common genes between ASD and catatonia for this chromosome.

Autistic Disorder↗

GAPs galore! A survey of putative Ras superfamily GTPase activating proteins in man and Drosophila.

Typical members of the Ras superfamily of small monomeric GTP-binding proteins function as regulators of diverse processes by cycling between biologically active GTP- and inactive GDP-bound conformations. Proteins that control this cycling include guanine nucleotide exchange factors or GEFs, which activate Ras superfamily members by catalyzing GTP for GDP exchange, and GTPase activating proteins or GAPs, which accelerate the low intrinsic GTP hydrolysis rate of typical Ras superfamily members, thus causing their inactivation. Two among the latter class of proteins have been implicated in common genetic disorders associated with an increased cancer risk, neurofibromatosis-1, and tuberous sclerosis. To facilitate genetic analysis, I surveyed Drosophila and human sequence databases for genes predicting proteins related to GAPs for Ras superfamily members. Remarkably, close to 0.5% of genes in both species (173 human and 64 Drosophila genes) predict proteins related to GAPs for Arf, Rab, Ran, Rap, Ras, Rho, and Sar family GTPases. Information on these genes has been entered into a pair of relational databases, which can be used to identify evolutionary conserved proteins that are likely to serve basic biological functions, and which can be updated when definitive information on the coding potential of both genomes becomes available.

ADP-Ribosylation Factors↗

Epistasis and balanced polymorphism influencing complex trait variation.

Complex traits such as human disease, growth rate, or crop yield are polygenic, or determined by the contributions from numerous genes in a quantitative manner. Although progress has been made in identifying major quantitative trait loci (QTL), experimental constraints have limited our knowledge of small-effect QTL, which may be responsible for a large proportion of trait variation. Here, we identified and dissected a one-centimorgan chromosome interval in Arabidopsis thaliana without regard to its effect on growth rate, and examined the signature of historical sequence polymorphism among Arabidopsis accessions. We found that the interval contained two growth rate QTL within 210 kilobases. Both QTL showed epistasis; that is, their phenotypic effects depended on the genetic background. This amount of complexity in such a small area suggests a highly polygenic architecture of quantitative variation, much more than previously documented. One QTL was limited to a single gene. The gene in question displayed a nucleotide signature indicative of balancing selection, and its phenotypic effects are reversed depending on genetic background. If this region typifies many complex trait loci, then non-neutral epistatic polymorphism may be an important contributor to genetic variation in complex traits.

Arabidopsis↗

Encoded evidence: DNA in forensic analysis.

Sherlock Holmes said "it has long been an axiom of mine that the little things are infinitely the most important", but never imagined that such a little thing, the DNA molecule, could become perhaps the most powerful single tool in the multifaceted fight against crime. Twenty years after the development of DNA fingerprinting, forensic DNA analysis is key to the conviction or exoneration of suspects and the identification of victims of crimes, accidents and disasters, driving the development of innovative methods in molecular genetics, statistics and the use of massive intelligence databases.

DNA Fingerprinting↗

Comprehensive evaluation of ACMG/AMP-based variant classification tools.

MOTIVATION: The American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines represent the gold standard for clinical variant interpretation. Despite the widespread adoption of ACMG/AMP guidelines, a comprehensive comparison of the software tools designed to implement them has been lacking. This represents a significant gap, as clinicians require evidence-based guidance on which tools to use in their practice. RESULTS: We benchmarked four ACMG/AMP-based tools (Franklin, InterVar, TAPES, Genebe) selected from 22 tools, and compared their performance with LIRICAL, a top-performing phenotype-driven tool, using 151 expert-curated datasets from Mendelian disorders. Selection criteria included free availability, VCF compatibility, operational reliability, and not being disease-specific. Our evaluation framework assessed top-N accuracy (N&#x2009;=&#x2009;1, 5, 10, 20, 50), retention rates, precision, recall, F1 scores, and area under the curve (AUC). Statistical validation employed bootstrap confidence intervals (n&#x2009;=&#x2009;1000) and Friedman tests. LIRICAL (68.21%) and Franklin (61.59%) demonstrated superior top-10 variant prioritization accuracy in Mendelian disorders, significantly outperforming other tools (P&#x2009;=&#x2009;.0000). Results demonstrate that tools with advanced phenotypic integration significantly outperform those relying primarily on genomic features. AVAILABILITY AND IMPLEMENTATION: All data and source code required to reproduce the findings of this study are openly available in the Code Ocean repository at https://doi.org/10.24433/CO.6562438.v1.

Software↗

PlantPan: A comprehensive multi-species plant pan-genome database.

The pan-genome represents the complete genomic diversity of specific species, serving as a valuable resource for studying species evolution, crop domestication, and guiding crop breeding and improvement. While there are several single-species-specific plant pan-genome databases, the availability of multi-species pan-genome databases is limited. Additionally, variations in methods and data types used for plant pan-genome analysis across different databases hinder the comparison and integration of pan-genome information from various projects at multi-species or single-species levels. To tackle this challenge, we introduce PlantPan, a comprehensive database housing the results of pan-genome analysis for 195 genomes from 11 plant species. PlantPan aims to provide extensive information, including gene-centric and sequence-centric pan-genome information, graph-based pan-genome, pan-genome openness profiles, gene functions and its variation characteristics, homologous genes, and gene clusters across different species. Statistically, PlantPan incorporates 9&#x2009;163&#x2009;011 genes, 694&#x2009;191 gene clusters, 526&#x2009;973&#x2009;370 genome variations, and 1&#x2009;616&#x2009;089 non-redundant genome variation groups at the species level, 33&#x2009;455,098 genome synteny, and 177&#x2009;827 non-redundant genome synteny groups at the species level. Regarding functional genes, PlantPan contains 5&#x2009;222&#x2009;720 genes related to transcription factors, 395&#x2009;247 literature-reported resistance genes, 455&#x2009;748 predicted microbial/disease resistance genes, and 1&#x2009;612&#x2009;112 genes related to molecular pathways. In summary, PlantPan is a vital platform for advancing the application of pan-genomes in molecular breeding for crops and evolutionary research for plants.

Genome, Plant↗

Molecular characterization and chromosomal distribution of Galileo, Kepler and Newton, three foldback transposable elements of the Drosophila buzzatii species complex.

Galileo is a foldback transposable element that has been implicated in the generation of two polymorphic chromosomal inversions in Drosophila buzzatii. Analysis of the inversion breakpoints led to the discovery of two additional elements, called Kepler and Newton, sharing sequence and structural similarities with Galileo. Here, we describe in detail the molecular structure of these three elements, on the basis of the 13 copies found at the inversion breakpoints plus 10 additional copies isolated during this work. Similarly to the foldback elements described in other organisms, these elements have long inverted terminal repeats, which in the case of Galileo possess a complex structure and display a high degree of internal variability between copies. A phylogenetic tree built with their shared sequences shows that the three elements are closely related and diverged approximately 10 million years ago. We have also analyzed the abundance and chromosomal distribution of these elements in D. buzzatii and other species of the repleta group by Southern analysis and in situ hybridization. Overall, the results suggest that these foldback elements are present in all the buzzatti complex species and may have played an important role in shaping their genomes. In addition, we show that recombination rate is the main factor determining the chromosomal distribution of these elements.

Animals↗

Microfluidic arrays in genetic analysis.

The goal of genetic analysis is to discover genetic markers that are informative for providing high confidence, positive predictive value in managing phenotypic outcomes. Primary consensus sequence data, genetic polymorphism databases and associated phenotype data are rapidly making genetic analysis more useful. Therefore, genetic analysis applications are gradually becoming more mainstream. The diversity and complexity of genetic analysis currently requires an array of analytical techniques, instrument platforms and software to support all the steps from data acquisition to interpretation. As supporting research technologies mature, they are incorporating increasing levels of automation, system integration and miniaturization. Microfluidic arrays are positioned to play a key role in routine genetic analysis, particularly as they begin to appear in more fully integrated analytical platforms.

Animals↗

MENDB: a database of polymorphic loci from natural populations.

MENDB is an online database for genetic markers determined to be polymorphic in natural populations. The database contains primer sequences and conditions for PCR, taxonomic information, and links to GenBank records, as well as basic statistics on the level of polymorphism for the surveyed populations/individuals.

Database Management Systems↗