Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “genetic databases”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Clinical and genetic heterogeneity in autosomal dominant cataract.

AIMS: To determine the different morphologies of autosomal dominant cataract (ADC), assess the intra- and interfamilial variation in cataract morphology, and undertake a genetic linkage study to identify loci for genes causing ADC and detect the underlying mutation. METHODS: Patients were recruited from the ocular genetic database at Moorfields Eye Hospital. All individuals underwent an eye examination with particular attention to the lens including anterior segment photography where possible. Blood samples were taken for DNA extraction and genetic linkage analysis was carried out using polymorphic microsatellite markers. RESULTS: 292 individuals from 16 large pedigrees with ADC were examined, of whom 161 were found to be affected. The cataract phenotypes could all be described as one of the eight following morphologies-anterior polar, posterior polar, nuclear, lamellar, coralliform, blue dot (cerulean), cortical, and pulverulent. The phenotypes varied in severity but the morphology was consistent within each pedigree, except in those affected by the pulverulent cataract, which showed considerable intrafamilial variation. Positive linkage was obtained in five families; in two families linkage was demonstrated to new loci and in three families linkage was demonstrated to previously described loci. In one of the families the underlying mutation was isolated. Exclusion data were obtained on five families. CONCLUSIONS: Although there is considerable clinical heterogeneity in ADC, the phenotype is usually consistent within families. There is extensive genetic heterogeneity and specific cataract phenotypes appear to be associated with mutations at more than one chromosome locus. In cases where the genetic mutation has been identified the molecular biology and clinical phenotype are closely associated.

Adolescent↗

Using information theory to search for co-evolving residues in proteins.

MOTIVATION: Some functionally important protein residues are easily detected since they correspond to conserved columns in a multiple sequence alignment (MSA). However important residues may also mutate, with compensatory mutations occurring elsewhere in the protein, which serve to preserve or restore functionality. It is difficult to distinguish these co-evolving sites from other non-conserved sites. RESULTS: We used Mutual Information (MI) to identify co-evolving positions. Using in silico evolved MSAs, we examined the effects of the number of sequences, the size of amino acid alphabet and the mutation rate on two sources of background MI: finite sample size effects and phylogenetic influence. We then assessed the performance of various normalizations of MI in enhancing detection of co-evolving positions and found that normalization by the pair entropy was optimal. Real protein alignments were analyzed and co-evolving isolated pairs were often found to be in contact with each other. AVAILABILITY: All data and program files can be found at http://www.biochem.uwo.ca/cgi-bin/CDD/index.cgi

Amino Acid Sequence↗

Large scale hierarchical clustering of protein sequences.

BACKGROUND: Searching a biological sequence database with a query sequence looking for homologues has become a routine operation in computational biology. In spite of the high degree of sophistication of currently available search routines it is still virtually impossible to identify quickly and clearly a group of sequences that a given query sequence belongs to. RESULTS: We report on our developments in grouping all known protein sequences hierarchically into superfamily and family clusters. Our graph-based algorithms take into account the topology of the sequence space induced by the data itself to construct a biologically meaningful partitioning. We have applied our clustering procedures to a non-redundant set of about 1,000,000 sequences resulting in a hierarchical clustering which is being made available for querying and browsing at http://systers.molgen.mpg.de/. CONCLUSIONS: Comparisons with other widely used clustering methods on various data sets show the abilities and strengths of our clustering methods in producing a biologically meaningful grouping of protein sequences.

Algorithms↗

Sample size, library composition, and genotypic diversity among natural populations of Escherichia coli from different animals influence accuracy of determining sources of fecal pollution.

A horizontal, fluorophore-enhanced, repetitive extragenic palindromic-PCR (rep-PCR) DNA fingerprinting technique (HFERP) was developed and evaluated as a means to differentiate human from animal sources of Escherichia coli. Box A1R primers and PCR were used to generate 2,466 rep-PCR and 1,531 HFERP DNA fingerprints from E. coli strains isolated from fecal material from known human and 12 animal sources: dogs, cats, horses, deer, geese, ducks, chickens, turkeys, cows, pigs, goats, and sheep. HFERP DNA fingerprinting reduced within-gel grouping of DNA fingerprints and improved alignment of DNA fingerprints between gels, relative to that achieved using rep-PCR DNA fingerprinting. Jackknife analysis of the complete rep-PCR DNA fingerprint library, done using Pearson's product-moment correlation coefficient, indicated that animal and human isolates were assigned to the correct source groups with an 82.2% average rate of correct classification. However, when only unique isolates were examined, isolates from a single animal having a unique DNA fingerprint, Jackknife analysis showed that isolates were assigned to the correct source groups with a 60.5% average rate of correct classification. The percentages of correctly classified isolates were about 15 and 17% greater for rep-PCR and HFERP, respectively, when analyses were done using the curve-based Pearson's product-moment correlation coefficient, rather than the band-based Jaccard algorithm. Rarefaction analysis indicated that, despite the relatively large size of the known-source database, genetic diversity in E. coli was very great and is most likely accounting for our inability to correctly classify many environmental E. coli isolates. Our data indicate that removal of duplicate genotypes within DNA fingerprint libraries, increased database size, proper methods of statistical analysis, and correct alignment of band data within and between gels improve the accuracy of microbial source tracking methods.

Animals↗

[Expression analysis of mouse homologous proteins with human aldose reductase like-1].

OBJECTIVE: To detect expression of mouse ARL-1 homologous proteins in mouse tissues, and analyze homology, genetic distance and phylogenetic relationship between human aldose reductase like-1 (ARL-1) and mouse homologous proteins. METHODS: Homology of mouse ARL-1 homologous proteins with human ARL-1 was analyzed by software Clustal X 1.8 using GenBank and Swiss-Prot database; genetic distance and phylogenetic relationship between mouse ARL-1 homologous proteins and human ARL-1 were analyzed by software Mega 2.0; mouse tissues were detected by Western blotting using polyclonal antibodies against ARL-1 protein from domestic rabbits. RESULTS: The amino acid sequence of human ARL-1 was 83%, 82%, 81%, 79%, 70%, 51%, 50% and 45% identical to that of the Chinese hamster ovary reductase (CHO-Red), the mouse fibroblast growth factor-regulated protein (FR-1), rat aldose reductase-like protein (rARLP), the mouse vas deferens protein (MVDP), rat lens aldose reductase (LeAR), delta4-3-ketosteroid-5beta-reductase (5beta-Red), rat aldo-keto reductase protein c (RaK-c) and 3alpha-hydroxysteroid dehydrogenase (3alpha-HSD). Among all the mouse ARL-1 homologous proteins, the genetic distance between CHO-Red and human ARL was the shortest (18.0%, P = 0.023), next was FR-1 (19.1%, P=0.023) and rARLP (19.9%, P = 0.025). From the phylogenetic tree, the protein whose relationship with human ARL-1 was the closest with CHO-Red, next was mouse FR-1, rARLP, MVDP and LeAR. Homologous proteins were found in mouse tissues including vas deferens, testis, bladder and uterus by Western blotting using polyclonal antibodies against ARL-1 protein from domestic rabbits. CONCLUSIONS: CHO-Red has the highest homology, the shortest genetic distance and the closest relationship with human ARL-1, next is FR-1, rARLP, MVDP. The major distribution of mouse ARL-1 homologous proteins is in vas deferens, testis, bladder and uterus, deducing they might be CHO-Red, FR-1, rARLP or MVDP

ADP-Ribosylation Factors↗

[Genetic polymorphism of 6 short-tandem repeat loci in Miao minority group of Rongshui county in Guangxi province].

OBJECTIVE: To investigate the distributions of six short-tandem repeat (STR) loci, namely D7S820, D13S317, D16S539, HUMCSF1PO, HUMTPOX and HUMTH01, in Miao minority group at Rongshui county in Guangxi province and construct the relevant genetic database. METHODS: Sodium-citrated blood specimens were collected from 208 healthy unrelated Miao individuals in Rongshui county. The DNAs from the specimens were extracted with phenol-chloroform method; AmplFSTR Identifier PCR Amplification Kit was used to amplify the extracted DNAs, and 3100 Genetic Analyzer was used to analyze and screen the amplified products. RESULTS: In this study, 7, 8, 6, 7, 5, 7 alleles were observed at the 6 STR loci respectively. The expected distribution of genotype accorded with Hardy-Weinbery equilibrium. The total discrimination power, cumulative paternity exclusion power and total polymorphism information were 0.999995, 0.9959 and 0.9987 respectively. CONCLUSION: The results demonstrate that these 6 STR loci are of high polymorphism and hereditary stability and are in accord with Mendel's law. The data obtained are valuable in population genetics research, forensic application, and individual identifications.

Adolescent↗

Isolation, characterization, and molecular identification of bacteria from the red imported fire ant (Solenopsis invicta) midgut.

Bacteria were isolated and cultured from the red imported fire ant (Solenopsis invicta) midgut. The small-subunit ribosomal RNA gene, (16s rRNA gene, approximately 1500 bp) was amplified from bacterial genomic DNA using the polymerase chain reaction and consensus sequence primers. Restriction fragment length polymorphism analysis revealed 10 unique profiles, indicating that at least 10 different bacteria are present in red imported fire ant midguts. The 16s rRNA gene sequence was determined for these isolates and queried against the NCBI genetic database. The results identified all isolates to at least the genus level. Antibiotic resistance profiles and biochemical activities were also determined for these species. This work provides the basis for a wider characterization of bacterial distributions in fire ant colonies and provides strains suitable for genetic manipulation to develop novel methods of fire ant control.

Animals↗

[Polymorphism of nine STR locus in Nu population from Yunnan Province].

In this study,blood samples were randomly drawn from 84 unrelated Nu individuals. The polymorphism of nine STR loci and Amelogenin locus were determined by DNA GeneScan. The genetic database on the distribution of gene frequency on the nine STR loci was established, statistical results showed that the genotype distributions were in agreement with Hardy-Weinberg equation. Compared with other population,the results in our study were of great value in human DNA genetic data instant method with the characteristics of precision and sensitivity.

English Abstract↗

[Distribution of STR locus DXS8027 polymorphism in the Han population in Qinba mountain areas].

To investigate the genetic polymorphism at DXS8027 STR locus for Han population in Qinba Mountain Areas and construct a preliminary population genetics database for this locus. We collected 550 venous blood samples from unrelated Han individuals of the Qinba Mountain Areas, Blood was treated with anticoagulant EDTA, and genomic DNA was obtained by phenol-chloroform extraction. Target fragments for the DXS8027 were amplified by PCR, separated on 8% non-denaturing polyacrylamide gel. electrophoresis and stained with 0.1% silver nitrate (AgNO3). Nine alleles were observed among the total samples, and the allele frequencies were in accordance with Hardy-Weinberg equilibrium (P>0.05) with relatively high heterozygosity (Het=0.7968). The results showed that there was a significant difference in the distribution of DXS8027 allele frequency between males and females (Chi 2=30.242, P<0.01), but the difference was not significant in the same gender groups of the two areas (Chi 2=4.703, P >0.05; Chi 2=14.952, P >0.05), or between the two areas regardless of gender (Chi 2=15.2, P >0.05). There was a very significant difference in DXS8027 allele frequency in populations of the Qinba Mountain Areas and the European populations, (Chi 2=37.572, P<0.01). It provides database to further study this STR locus in different population.

Asian People↗

Novel COL7A1 mutations in dystrophic forms of epidermolysis bullosa.

Mutations in the type VII collagen gene (COL7A1) have been shown to underlie different variants of dystrophic epidermolysis bullosa (DEB). Examination of the genetic database indicates that most of the mutations are family specific, with few recurrent mutations. To facilitate further refinement of genotype/phenotype correlations in DEB, we have examined a cohort of nine families with DEB (seven recessively and two dominantly inherited) by a mutation detection strategy based on polymerase chain reaction amplification of COL7A1 genomic sequences, followed by heteroduplex scanning and direct nucleotide sequencing. The results revealed 16 allelic mutations, 11 of them being novel, previously unpublished. The genetic information was also used for prenatal testing in a family at risk for recurrence of a severe, Hallopeau-Siemens type of RDEB. These data contribute to the expanding database of COL7A1 mutations in DEB.

Adult↗

Molecular models of NS3 protease variants of the Hepatitis C virus.

BACKGROUND: Hepatitis C virus (HCV) currently infects approximately three percent of the world population. In view of the lack of vaccines against HCV, there is an urgent need for an efficient treatment of the disease by an effective antiviral drug. Rational drug design has not been the primary way for discovering major therapeutics. Nevertheless, there are reports of success in the development of inhibitor using a structure-based approach. One of the possible targets for drug development against HCV is the NS3 protease variants. Based on the three-dimensional structure of these variants we expect to identify new NS3 protease inhibitors. In order to speed up the modeling process all NS3 protease variant models were generated in a Beowulf cluster. The potential of the structural bioinformatics for development of new antiviral drugs is discussed. RESULTS: The atomic coordinates of crystallographic structure 1CU1 and 1DY9 were used as starting model for modeling of the NS3 protease variant structures. The NS3 protease variant structures are composed of six subdomains, which occur in sequence along the polypeptide chain. The protease domain exhibits the dual beta-barrel fold that is common among members of the chymotrypsin serine protease family. The helicase domain contains two structurally related beta-alpha-beta subdomains and a third subdomain of seven helices and three short beta strands. The latter domain is usually referred to as the helicase alpha-helical subdomain. The rmsd value of bond lengths and bond angles, the average G-factor and Verify 3D values are presented for NS3 protease variant structures. CONCLUSIONS: This project increases the certainty that homology modeling is an useful tool in structural biology and that it can be very valuable in annotating genome sequence information and contributing to structural and functional genomics from virus. The structural models will be used to guide future efforts in the structure-based drug design of a new generation of NS3 protease variants inhibitors. All models in the database are publicly accessible via our interactive website, providing us with large amount of structural models for use in protein-ligand docking analysis.

Amino Acid Sequence↗

The adaptive evolution database (TAED).

BACKGROUND: The Master Catalog is a collection of evolutionary families, including multiple sequence alignments, phylogenetic trees and reconstructed ancestral sequences, for all protein-sequence modules encoded by genes in GenBank. It can therefore support large-scale genomic surveys, of which we present here The Adaptive Evolution Database (TAED). In TAED, potential examples of positive adaptation are identified by high values for the normalized ratio of nonsynonymous to synonymous nucleotide substitution rates (KA/KS values) on branches of an evolutionary tree between nodes representing reconstructed ancestral sequences. RESULTS: Evolutionary trees and reconstructed ancestral sequences were extracted from the Master Catalog for every subtree containing proteins from the Chordata only or the Embryophyta only. Branches with high KA/KS values were identified. These represent candidate episodes in the history of the protein family when the protein may have undergone positive selection, where the mutant form conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. An unexpectedly large number of families (between 10% and 20% of those families examined) were found to have at least one branch with high KA/KS values above arbitrarily chosen cut-offs (1 and 0.6). Most of these survived a robustness test and were collected into TAED. CONCLUSIONS: TAED is a raw resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study. It can be expanded to include other evolutionary information (for example changes in gene regulation or splicing) placed in a phylogenetic perspective.

Adaptation, Physiological↗

Who tangos with GOA?-Use of Gene Ontology Annotation (GOA) for biological interpretation of '-omics' data and for validation of automatic annotation tools.

The number of large-scale experimental datasets generated from high-throughput technologies has grown rapidly. Biological knowledge resources such as the Gene Ontology Annotation (GOA) database, which provides high-quality functional annotation to proteins within the UniProt Knowledgebase, can play an important role in the analysis of such data. The integration of GOA with analytical tools has proved to aid the clustering, annotation and biological interpretation of such large expression datasets. GOA is also useful in the development and validation of automated annotation tools, in particular text-mining systems. The increasing interest in GOA highlights the great potential of this freely available resource to assist both the biological research and bioinformatics communities.

Animals↗

Mycobacterium tuberculosis complex genetic diversity: mining the fourth international spoligotyping database (SpolDB4) for classification, population genetics and epidemiology.

BACKGROUND: The Direct Repeat locus of the Mycobacterium tuberculosis complex (MTC) is a member of the CRISPR (Clustered regularly interspaced short palindromic repeats) sequences family. Spoligotyping is the widely used PCR-based reverse-hybridization blotting technique that assays the genetic diversity of this locus and is useful both for clinical laboratory, molecular epidemiology, evolutionary and population genetics. It is easy, robust, cheap, and produces highly diverse portable numerical results, as the result of the combination of (1) Unique Events Polymorphism (UEP) (2) Insertion-Sequence-mediated genetic recombination. Genetic convergence, although rare, was also previously demonstrated. Three previous international spoligotype databases had partly revealed the global and local geographical structures of MTC bacilli populations, however, there was a need for the release of a new, more representative and extended, international spoligotyping database. RESULTS: The fourth international spoligotyping database, SpolDB4, describes 1939 shared-types (STs) representative of a total of 39,295 strains from 122 countries, which are tentatively classified into 62 clades/lineages using a mixed expert-based and bioinformatical approach. The SpolDB4 update adds 26 new potentially phylogeographically-specific MTC genotype families. It provides a clearer picture of the current MTC genomes diversity as well as on the relationships between the genetic attributes investigated (spoligotypes) and the infra-species classification and evolutionary history of the species. Indeed, an independent Naïve-Bayes mixture-model analysis has validated main of the previous supervised SpolDB3 classification results, confirming the usefulness of both supervised and unsupervised models as an approach to understand MTC population structure. Updated results on the epidemiological status of spoligotypes, as well as genetic prevalence maps on six main lineages are also shown. Our results suggests the existence of fine geographical genetic clines within MTC populations, that could mirror the passed and present Homo sapiens sapiens demographical and mycobacterial co-evolutionary history whose structure could be further reconstructed and modelled, thereby providing a large-scale conceptual framework of the global TB Epidemiologic Network. CONCLUSION: Our results broaden the knowledge of the global phylogeography of the MTC complex. SpolDB4 should be a very useful tool to better define the identity of a given MTC clinical isolate, and to better analyze the links between its current spreading and previous evolutionary history. The building and mining of extended MTC polymorphic genetic databases is in progress.

Computational Biology↗

The use of routinely collected computer data for research in primary care: opportunities and challenges.

INTRODUCTION: Routinely collected primary care data has underpinned research that has helped define primary care as a specialty. In the early years of the discipline, data were collected manually, but digital data collection now makes large volumes of data readily available. Primary care informatics is emerging as an academic discipline for the scientific study of how to harness these data. This paper reviews how data are stored in primary care computer systems; current use of large primary care research databases; and, the opportunities and challenges for using routinely collected primary care data in research. OPPORTUNITIES: (1) Growing volumes of routinely recorded data. (2) Improving data quality. (3) Technological progress enabling large datasets to be processed. (4) The potential to link clinical data in family practice with other data including genetic databases. (5) An established body of know-how within the international health informatics community. CHALLENGES: (1) Research methods for working with large primary care datasets are limited. (2) How to infer meaning from data. (3) Pace of change in medicine and technology. (4) Integrating systems where there is often no reliable unique identifier and between health (person-based records) and social care (care-based records-e.g. child protection). (5) Achieving appropriate levels of information security, confidentiality, and privacy. CONCLUSION: Routinely collected primary care computer data, aggregated into large databases, is used for audit, quality improvement, health service planning, epidemiological study and research. However, gaps exist in the literature about how to find relevant data, select appropriate research methods and ensure that the correct inferences are drawn.

Biomedical Research↗

A new pentaplex PCR system for forensic casework analysis.

In 1998 the Federal Criminal Police Office of Germany (BKA) established a central genetic database of offenders and suspects to facilitate comparisons with biological samples from future criminal offences. The five obligatory short tandem repeat (STR) loci in this database (TH01, SE33, vWA, FGA and D21S11) were co-amplified in a new PCR pentaplex analysing system together with the sex-specific locus amelogenin. Due to overlapping fragment sizes, amplification products were fluorescent dye-labelled with different colours, separated by electrophoresis and detected directly using the ABI PRISM 310 Genetic Analyzer. Reproducible and reliable results were obtained from as low as 125 pg template DNA, indicating high specificity and sensitivity of the assay. Environmental studies and enzymatic digest with DNase I revealed an excellent stability of the pentaplex system with typeable results even in cases of partially degraded DNA. Complete and reproducible DNA typing was possible in blood-stain mixtures with the minor component as low as 10%. Mean stutter peak intensities were analysed for all loci and ranged from 2.7 +/- 0.8% (TH01) to 10.6 +/- 1.6% (vWA) of the main signal intensity. Allele frequencies were determined in a North Bavarian population sample (n = 121). The combination of five systems resulted in a mean exclusion chance of 99.86% and a power of discrimination of 99.999996%. No deviation from Hardy-Weinberg equilibrium could be found.

Alleles↗

The forensic DNA implications of genetic differentiation between endogamous communities.

In many indigenous minority populations, and among migrants from Asian and African populations now resident in western Europe, North America and Australia, there is a strong tradition of endogamy and a preference for consanguineous unions. These marriage practices can result in F(ST) values greatly in excess of the maximum value (0.01) currently recommended for forensic DNA purposes under guidelines established by the National Research Council (NRC) of the USA. To examine the possible extent of deviation from this accepted norm, three co-resident Pakistani communities were studied using 10 autosomal dinucleotide markers and six tetranucleotide markers on the Y-chromosome. The mean population subdivision coefficient (FST) value was 0.13 for the autosomal loci, and Y-chromosome loci exhibited even stronger differentiation with unique alleles identified in all three communities. The data indicate that even when sub-populations are virtually indistinguishable in terms of anthropology, geography, ethnicity or culture, they may still exhibit major genetic differentiation. Where significant population stratification is known to exist, more detailed genetic databases should be developed for forensic DNA purposes, based on reference data from each of the appropriate sub-populations and not on random or combined samples.

Consanguinity↗

Proteometric study of ghrelin receptor function variations upon mutations using amino acid sequence autocorrelation vectors and genetic algorithm-based least square support vector machines.

Functional variations on the human ghrelin receptor upon mutations have been associated with a syndrome of short stature and obesity, of which the obesity appears to develop around puberty. In this work, we reported a proteometrics analysis of the constitutive and ghrelin-induced activities of wild-type and mutant ghrelin receptors using amino acid sequence autocorrelation (AASA) approach for protein structural information encoding. AASA vectors were calculated by measuring the autocorrelations at sequence lags ranging from 1 to 15 on the protein primary structure of 48 amino acid/residue properties selected from the AAindex database. Genetic algorithm-based multilinear regression analysis (GA-MRA) and genetic algorithm-based least square support vector machines (GA-LSSVM) were used for building linear and non-linear models of the receptor activity. A genetic optimized radial basis function (RBF) kernel yielded the optimum GA-LSSVM models describing 88% and 95% of the cross-validation variance for the constitutive and ghrelin-induced activities, respectively. AASA vectors in the optimum models mainly appeared weighted by hydrophobicity-related properties. However, differently to the constitutive activity, the ghrelin-induced activity was also highly dependent of the steric features of the receptor.

Algorithms↗