Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Genetic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Genetics of the immune response: identifying immune variation within the MHC and throughout the genome.

With the advent of modern genomic sequencing technology the ability to obtain new sequence data and to acquire allelic polymorphism data from a broad range of samples has become routine. In this regard, our investigations have started with the most polymorphic of genetic regions fundamental to the immune response in the major histocompatibility complex (MHC). Starting with the completed human MHC genomic sequence, we have developed a resource of methods and information that provide ready access to a large portion of human and nonhuman primate MHCs. This resource consists of a set of primer pairs or amplicons that can be used to isolate about 15% of the 4.0 Mb MHC. Essentially similar studies are now being carried out on a set of immune response loci to broaden the usefulness of the data and tools developed. A panel of 100 genes involved in the immune response have been targeted for single nucleotide polymorphism (SNP) discovery efforts that will analyze 120 Mb of sequence data for the presence of immune-related SNPs. The SNP data provided from the MHC and from the immune response panel has been adapted for use in studies of evolution, MHC disease associations, and clinical transplantation.

Animals↗

Genetic analysis of anterior posterior expression gradients in the developing mammalian forebrain.

Intrinsic regulatory factors play critical roles in early cortical patterning, including the development of the anteroposterior (A-P) axis. To identify genes that are differentially expressed along the A-P axis of the developing cerebral cortex, we analyzed gene expression in presumptive frontal, parietal, and occipital cerebral walls of E12.5 mouse using complementary DNA microarrays. We identified 106 genes, including expressed sequence tags (ESTs), expressed in an A-P gradient in the embryonic brain and screened 88 by in situ hybridization for confirmation. Central nervous system (CNS) expression patterns of many of these genes were previously unknown. Others, such as Sfrp1, CoupTF1, and FABP7, were expressed in a manner consistent with previous studies, providing independent confirmation. Two related transcription factors, previously not implicated in CNS development, Fhl1 and Fhl2, were observed to be enriched in posterior and anterior telencephalon, respectively. We studied patterning gradients in Fhl1 knockout mice but observed no changes in gene expression related to A-P regionalization in the Fhl1 knockout mice. These data provide an important set of new candidates for studies of cortical patterning and maturation.

Animals↗

Analysis of quantitative lipid traits in the genetics of NIDDM (GENNID) study.

Coronary heart disease (CHD) is the leading cause of death among individuals with type 2 diabetes. Dyslipidemia contributes significantly to CHD in diabetic patients, in whom lipid abnormalities include hypertriglyceridemia, low HDL cholesterol, and increased levels of small, dense LDL particles. To identify genes for lipid-related traits, we performed genome-wide linkage analyses for levels of triglycerides and HDL, LDL, and total cholesterol in Caucasian, Hispanic, and African-American families from the Genetics of NIDDM (GENNID) study. Most lipid traits showed significant estimates of heritability (P < 0.001) with the exception of triglycerides and the triglyceride/HDL ratio in African Americans. Variance components analysis identified linkage on chromosome 3p12.1-3q13.31 for the triglyceride/HDL ratio (logarithm of odds [LOD] = 3.36) and triglyceride (LOD = 3.27) in Caucasian families. Statistically significant evidence for linkage was identified for the triglyceride/HDL ratio (LOD = 2.45) on 11p in Hispanic families in a region that showed suggestive evidence for linkage (LOD = 2.26) for triglycerides in this population. In African Americans, the strongest evidence for linkage (LOD = 2.26) was found on 19p13.2-19q13.42 for total cholesterol. Our findings provide strong support for previous reports of linkage for lipid-related traits, suggesting the presence of genes on 3p12.1-3q13.31, 11p15.4-11p11.3, and 19p13.2-19q13.42 that may influence traits underlying lipid abnormalities associated with type 2 diabetes.

Adult↗

The Legume Information System (LIS): an integrated information resource for comparative legume biology.

The Legume Information System (LIS) (http://www.comparative-legumes.org), developed by the National Center for Genome Resources in cooperation with the USDA Agricultural Research Service (ARS), is a comparative legume resource that integrates genetic and molecular data from multiple legume species enabling cross-species genomic and transcript comparisons. The LIS virtual plant interface allows simplified and intuitive navigation of transcript data from Medicago truncatula, Lotus japonicus, Glycine max and Arabidopsis thaliana. Transcript libraries are represented as images of plant organs in different developmental stages, which are selected to query the analyzed and annotated data. Complex queries can be accomplished by adding modifiers, keywords and sequence names. The LIS also contains annotated genomic data featuring transcript alignments to validate gene predictions as well as motif and similarity analyses. The genomic browser supports comparative analysis via novel dynamic functional annotation comparisons. CMap, developed as part of the GMOD project (http://www.gmod.org/cmap/index.shtml), has been incorporated to support comparative analyses of community linkage and physical map data. LIS is being expanded to incorporate gene expression and biochemical pathways which will be seamlessly integrated forming a knowledge discovery framework.

Arabidopsis↗

Gene selection using support vector machines with non-convex penalty.

MOTIVATION: With the development of DNA microarray technology, scientists can now measure the expression levels of thousands of genes simultaneously in one single experiment. One current difficulty in interpreting microarray data comes from their innate nature of 'high-dimensional low sample size'. Therefore, robust and accurate gene selection methods are required to identify differentially expressed group of genes across different samples, e.g. between cancerous and normal cells. Successful gene selection will help to classify different cancer types, lead to a better understanding of genetic signatures in cancers and improve treatment strategies. Although gene selection and cancer classification are two closely related problems, most existing approaches handle them separately by selecting genes prior to classification. We provide a unified procedure for simultaneous gene selection and cancer classification, achieving high accuracy in both aspects. RESULTS: In this paper we develop a novel type of regularization in support vector machines (SVMs) to identify important genes for cancer classification. A special nonconvex penalty, called the smoothly clipped absolute deviation penalty, is imposed on the hinge loss function in the SVM. By systematically thresholding small estimates to zeros, the new procedure eliminates redundant genes automatically and yields a compact and accurate classifier. A successive quadratic algorithm is proposed to convert the non-differentiable and non-convex optimization problem into easily solved linear equation systems. The method is applied to two real datasets and has produced very promising results. AVAILABILITY: MATLAB codes are available upon request from the authors.

Algorithms↗

Genetic test bed for feature selection.

MOTIVATION: Given a large set of potential features, such as the set of all gene-expression values from a microarray, it is necessary to find a small subset with which to classify. The task of finding an optimal feature set of a given size is inherently combinatoric because to assure optimality all feature sets of a given size must be checked. Thus, numerous suboptimal feature-selection algorithms have been proposed. There are strong impediments to evaluate feature-selection algorithms using real data when data are limited, a common situation in genetic classification. The difficulty is compound. First, there are no class-conditional distributions from which to draw data points, only a single small labeled sample. Second, there are no test data with which to estimate the feature-set errors, and one must depend on a training-data-based error estimator. Finally, there is no optimal feature set with which to compare the feature sets found by the algorithms. RESULTS: This paper describes a genetic test bed for the evaluation of feature-selection algorithms. It begins with a large biological feature-label dataset that is used as an empirical distribution and, using massively parallel computation, finds the top feature sets of various sizes based on a given sample size and classification rule. The user can draw random samples from the data, apply a proposed algorithm, and evaluate the proficiency of the proposed algorithm via three different measures (code provided). A key feature of the test bed is that, once a dataset is input, a single command creates the entire test bed relative to the dataset. The particular dataset used for the first version of the test bed comes from a microarray-based classification study that analyzes a large number of microarrays, prepared with RNA from breast tumor samples from each of 295 patients. AVAILABILITY: The software and supplementary material are available at http://public.tgen.org/tgen-cb/support/testbed/ CONTACT: edward@ece.tamu.edu.

Algorithms↗

The molecular genetics of Marfan syndrome and related disorders.

Marfan syndrome (MFS), a relatively common autosomal dominant hereditary disorder of connective tissue with prominent manifestations in the skeletal, ocular, and cardiovascular systems, is caused by mutations in the gene for fibrillin-1 (FBN1). The leading cause of premature death in untreated individuals with MFS is acute aortic dissection, which often follows a period of progressive dilatation of the ascending aorta. Recent research on the molecular physiology of fibrillin and the pathophysiology of MFS and related disorders has changed our understanding of this disorder by demonstrating changes in growth factor signalling and in matrix-cell interactions. The purpose of this review is to provide a comprehensive overview of recent advances in the molecular biology of fibrillin and fibrillin-rich microfibrils. Mutations in FBN1 and other genes found in MFS and related disorders will be discussed, and novel concepts concerning the complex and multiple mechanisms of the pathogenesis of MFS will be explained.

Activin Receptors, Type I↗

Using ancestry-informative markers to define populations and detect population stratification.

A serious problem with case-control studies is that population subdivision, recent admixture and sampling variance can lead to spurious associations between a phenotype and a marker locus, or indeed may mask true associations. This is also a concern in therapeutics since drug response may differ by ethnicity. Population stratification can occur if cases and controls have different frequencies of ethnic groups or in admixed populations, different fractions of ancestry, and when phenotypes of interest such as disease, drug response or drug metabolism, also differ between ethnic groups. Although most genetic variation is inter-individual, there is also significant inter-ethnic variation. The International HapMap Project has provided allele frequencies for approximately three million single nucleotide polymorphisms (SNPs) in Africans, Europeans and East Asians. SNP variation is greatest in Africans. Statistical methods for the detection and correction of population stratification, principally Structured Association and Genomic Control, have recently become freely available. These methods use marker loci spread throughout the genome that are unlinked to the candidate locus to estimate the ancestry of individuals within a sample, and to test for and adjust the ethnic matching of cases and controls. To date, few case-control association studies have incorporated testing for population stratification. This paper will focus on the debate about the quantity and methods for selection of highly informative marker loci required to characterize populations that vary in substructure or the degree of admixture, and will discuss how these theoretically desirable approaches can be effectively put into practice.

Case-Control Studies↗

Why highly expressed proteins evolve slowly.

Much recent work has explored molecular and population-genetic constraints on the rate of protein sequence evolution. The best predictor of evolutionary rate is expression level, for reasons that have remained unexplained. Here, we hypothesize that selection to reduce the burden of protein misfolding will favor protein sequences with increased robustness to translational missense errors. Pressure for translational robustness increases with expression level and constrains sequence evolution. Using several sequenced yeast genomes, global expression and protein abundance data, and sets of paralogs traceable to an ancient whole-genome duplication in yeast, we rule out several confounding effects and show that expression level explains roughly half the variation in Saccharomyces cerevisiae protein evolutionary rates. We examine causes for expression's dominant role and find that genome-wide tests favor the translational robustness explanation over existing hypotheses that invoke constraints on function or translational efficiency. Our results suggest that proteins evolve at rates largely unrelated to their functions and can explain why highly expressed proteins evolve slowly across the tree of life.

Computational Biology↗

Genome-wide tagging SNPs with entropy-based Monte Carlo method.

The number of common single nucleotide polymorphisms (SNPs) in the human genome is estimated to be around 3-6 million. It is highly anticipated that the study of SNPs will help provide a means for elucidating the genetic component of complex diseases and variable drug responses. High-throughput technologies such as oligonucleotide arrays have produced enormous amount of SNP data, which creates great challenges in genome-wide disease linkage and association studies. In this paper, we present an adaptation of the cross entropy (CE) method and propose an iterative CE Monte Carlo (CEMC) algorithm for tagging SNP selection. This differs from most of SNP selection algorithms in the literature in that our method is independent of the notion of haplotype block. Thus, the method is applicable to whole genome SNP selection without prior knowledge of block boundaries. We applied this block-free algorithm to three large datasets (two simulated and one real) that are in the order of thousands of SNPs. The successful applications to these large scale datasets demonstrate that CEMC is computationally feasible for whole genome SNP selection. Furthermore, the results show that CEMC is significantly better than random selection, and it also outperformed another block-free selection algorithm for the dataset considered.

Algorithms↗

A high-throughput Arabidopsis reverse genetics system.

A collection of Arabidopsis lines with T-DNA insertions in known sites was generated to increase the efficiency of functional genomics. A high-throughput modified thermal asymmetric interlaced (TAIL)-PCR protocol was developed and used to amplify DNA fragments flanking the T-DNA left borders from approximately 100000 transformed lines. A total of 85108 TAIL-PCR products from 52964 T-DNA lines were sequenced and compared with the Arabidopsis genome to determine the positions of T-DNAs in each line. Predicted T-DNA insertion sites, when mapped, showed a bias against predicted coding sequences. Predicted insertion mutations in genes of interest can be identified using Arabidopsis Gene Index name searches or by BLAST (Basic Local Alignment Search Tool) search. Insertions can be confirmed by simple PCR assays on individual lines. Predicted insertions were confirmed in 257 of 340 lines tested (76%). This resource has been named SAIL (Syngenta Arabidopsis Insertion Library) and is available to the scientific community at www.tmri.org.

Agrobacterium tumefaciens↗

Significance of the genetic relationships deduced from partial nucleotide sequencing of infectious bursal disease virus genome segments A or B.

The rapid genomic characterization of infectious bursal disease virus (IBDV) requires determining which partial nucleotide (nt) sequences derived from IBDV segments A or B would produce phylogenetic information as significant as sequencing the whole corresponding segments. Long nt coding sequences of 27 IBDV segments A (aa 20-991) and 21 segments B (aa 7-stop codon) were retrieved from databanks and used to compute reference phylogenetic trees using Neighbor Joining (NJ) and Parsimony (P): clusters appearing in the NJ and P reference trees with a bootstrap value greater than 80% were considered as significant (Whole Segment Clusters, WSC). The sequences were then cut into overlapping regions. These were used to compute phylogenetic trees which were compared with reference ones. Of the partial sequences, the VP2 gene best represented IBDV segment A (10 out of 13 WSC were conserved), and the 5' two thirds of segment B best represented segment B (5 to 6 conserved WSC out of 6). Implementation of the Plato programme finally demonstrated that the region encoding VP2 variable domain (vVP2, segment A) is the only region of IBDV genome with a significantly different evolution rate, which result is consistent with vVP2 being subjected to a high selection pressure.

Databases, Genetic↗

Estimating effective population size or mutation rate with microsatellites.

Microsatellites are short tandem repeats that are widely dispersed among eukaryotic genomes. Many of them are highly polymorphic; they have been used widely in genetic studies. Statistical properties of all measures of genetic variation at microsatellites critically depend upon the composite parameter theta = 4Nmicro, where N is the effective population size and micro is mutation rate per locus per generation. Since mutation leads to expansion or contraction of a repeat number in a stepwise fashion, the stepwise mutation model has been widely used to study the dynamics of these loci. We developed an estimator of theta, theta; (F), on the basis of sample homozygosity under the single-step stepwise mutation model. The estimator is unbiased and is much more efficient than the variance-based estimator under the single-step stepwise mutation model. It also has smaller bias and mean square error (MSE) than the variance-based estimator when the mutation follows the multistep generalized stepwise mutation model. Compared with the maximum-likelihood estimator theta; (L) by, theta; (F) has less bias and smaller MSE in general. theta; (L) has a slight advantage when theta is small, but in such a situation the bias in theta; (L) may be more of a concern.

Alleles↗

Evaluation of biobank constitution and use: multicentre analysis in France and propositions for formalising the activities of research ethics committees.

Biobanks are collections of biological material and related files gathered and stored for clinical or research purposes. Here, we investigated the questions raised during the evaluation of biobanks by biomedical Research Ethics Committees (RECs), particularly in the context of genetic research. We sent a questionnaire to all RECs in France to survey their concerns and the ethical criteria used when evaluating research involving the storage of biological samples. Most of the RECs think that they should be consulted to evaluate the constitution of biobanks. The proportion of RECs of this opinion depended on whether the biobank is being constituted in the absence of an associated research project (initially created for clinical purposes or for undefined research) (14/28), whether the biobank is being constituted for research use (21/28) or whether an existing research biobank is being re-used (19/28). Views diverged concerning the way ethics principles are applied, showing that REC evaluations of biobanks might be formalised at each of the following steps: constitution, use and re-use. In this paper, we suggest concrete elements that could be integrated into the application of the new French law concerning the protection of the human beings participating in research as well as into international recommendations.

Biomedical Research↗

Health implications of alpha1-antitrypsin deficiency in Sub-Sahara African countries and their emigrants in Europe and the New World.

PURPOSE: To determine the frequencies of the protease inhibitor (PI) deficiency alleles of alpha1-antitrypsin deficiency (AAT Deficiency) in indigenous populations in 12 countries in Sub-Sahara Africa because of their potential impact on the health in these populations with regard to the high risk for development of liver and lung disease. In addition, to discuss the unique susceptibility of these populations and emigrants to Europe and the New World to the adverse health effects associated with exposure to environmental microbes, chemicals, and particulates. METHODS: Detailed statistical analysis of the 24 control cohort databases from genetic epidemiological studies by others were used to estimate the allele frequencies and prevalence for the two most common deficiency alleles PIS and PIZ and to estimate the numbers at risk in each of the local Sub-Sahara populations as well as those who have emigrated from these countries to Europe and the New World. RESULTS: The present study has provided evidence for the presence of both PIS and PIZ in the general populations of Nigeria, Republic of South Africa, and Somalia, the PIS allele in Angola, Botswana, Cameroon, Mozambique, Namibia, and the Republic of Congo, and only the PIZ allele in Mali. CONCLUSION: AAT Deficiency is found in both the Black and "Colored" populations in many of the Sub-Sahara countries in Africa, providing evidence for the presence of AAT Deficiency in such populations in Europe and in the New World. Such populations should be screened for AAT Deficiency and made aware of their unique susceptibility to exposure to chemical and particulate agents in the environment.

Africa, Northern↗

A new integrated genetic linkage map of the soybean.

A total of 391 simple sequence repeat (SSR) markers designed from genomic DNA libraries, 24 derived from existing GenBank genes or ESTs, and five derived from bacterial artificial chromosome (BAC) end sequences were developed. In contrast to SSRs derived from EST sequences, those derived from genomic libraries were a superior source of polymorphic markers, given that the mean number of tandem repeats in the former was significantly less than that of the latter ( P<0.01). The 420 newly developed SSRs were mapped in one or more of five soybean mapping populations: "Minsoy" x "Noir 1", "Minsoy" x "Archer", "Archer" x "Noir 1", "Clark" x "Harosoy", and A81-356022 x PI468916. The JoinMap software package was used to combine the five maps into an integrated genetic map spanning 2,523.6 cM of Kosambi map distance across 20 linkage groups that contained 1,849 markers, including 1,015 SSRs, 709 RFLPs, 73 RAPDs, 24 classical traits, six AFLPs, ten isozymes, and 12 others. The number of new SSR markers added to each linkage group ranged from 12 to 29. In the integrated map, the ratio of SSR marker number to linkage group map distance did not differ among 18 of the 20 linkage groups; however, the SSRs were not uniformly spaced over a linkage group, clusters of SSRs with very limited recombination were frequently present. These clusters of SSRs may be indicative of gene-rich regions of soybean, as has been suggested by a number of recent studies, indicating the significant association of genes and SSRs. Development of SSR markers from map-referenced BAC clones was a very effective means of targeting markers to marker-scarce positions in the genome.

Chromosome Mapping↗

The modular nature of genetic diseases.

Evidence from many sources suggests that similar phenotypes are begotten by functionally related genes. This is most obvious in the case of genetically heterogeneous diseases such as Fanconi anemia, Bardet-Biedl or Usher syndrome, where the various genes work together in a single biological module. Such modules can be a multiprotein complex, a pathway, or a single cellular or subcellular organelle. This observation suggests a number of hypotheses about the human phenome that are now beginning to be explored. First, there is now good evidence from bioinformatic analyses that human genetic diseases can be clustered on the basis of their phenotypic similarities and that such a clustering represents true biological relationships of the genes involved. Second, one may use such phenotypic similarity to predict and then test for the contribution of apparently unrelated genes to the same functional module. This concept is now being systematically tested for several diseases. Most recently, a systematic yeast two-hybrid screen of all known genes for inherited ataxias indicated that they all form part of a single extended protein-protein interaction network. Third, one can use bioinformatics to make predictions about new genes for diseases that form part of the same phenotype cluster. This is done by starting from the known disease genes and then searching for genes that share one or more functional attributes such as gene expression pattern, coevolution, or gene ontology. Ultimately, one may expect that a modular view of disease genes should help the rapid identification of additional disease genes for multifactorial diseases once the first few contributing genes (or environmental factors) have been reliably identified.

Computational Biology↗

Candidate genes for nicotine dependence via linkage, epistasis, and bioinformatics.

Many smoking-related phenotypes are substantially heritable. One genome scan of nicotine dependence (ND) has been published and several others are in progress and should be completed in the next 5 years. The goal of this hypothesis-generating study was two-fold. First, we present further analyses of our genome scan data for ND published by Straub et al. [1999: Mol Psychiatry 4:129-144] (PMID: 10208445). Second, we used the method described by Cox et al. [1999: Nat Genet 21:213-215] (PMID: 9988276) to search for epistatic loci across the markers used in the genome scan. The overall results of the genome scan nearly reached the rigorous Lander and Kruglyak [1995: Nat Genet 11:241-247] criteria for "significant" linkage with the best findings on chromosomes 10 and 2. We then looked for correspondence between genes located in the 10 regions implicated in affected sibling pair (ASP) and epistatic linkage analyses with a list of genes suggested by microarray studies of experimental nicotine exposure and candidate genes from the literature. We found correspondence between linkage and microarray/candidate gene studies for genes involved with the mitogen-activated protein kinase (MAPK) signaling system, nuclear factor kappa B (NFKB) complex, neuropeptide Y (NPY) neurotransmission, a nicotinic receptor subunit (CHRNA2), the vesicular monoamine transporter (SLC18A2), genes in pathways implicated in human anxiety (HTR7, TDO2, and the endozepine-related protein precursor, DKFZP434A2417), and the micro 1-opioid receptor (OPRM1). Although the hypotheses resulting from these linkage and bioinformatic analyses are plausible and intriguing, their ultimate worth depends on replication in additional linkage samples and in future experimental studies.

Chromosome Mapping↗